Change in structure, with Minto and without
Each cell is one judged document: its structure score with Minto minus the same model’s score without it. The most structure can score is 8 points.
Claude Haiku 4.5 ran two audit cases in this run. On a set of headings its structure went from 2 out of 8 to 6. On a set of monthly reports it went from 6 down to 2. Same model, same mode, same skill, opposite results. That is why this evaluation is a map of cases rather than a single figure.
Run v1.7.0-final · 11 cases across 4 modes · 8 models · 172 blind verdicts over 86 pairs
Measure
Each cell is one judged document: its structure score with Minto minus the same model’s score without it. The most structure can score is 8 points.
Each cell is one judged document: its writing-quality score with Minto minus the same model’s score without it. The most quality can score is 10 points.
Reading the shading
One document per cell. Each cell is a single judged pair, not an average. A cell of plus or minus one is noise; the row means, over eight models, are the steadier number, and the modes are steadier still.
172 verdicts, 86 pairs. Every pair is judged twice, once without the skill and once with it, which is where the landing page’s 172 comes from. The two empty cells are qwen3.8-27b on the two longest cases; it timed out in both arms, so there is nothing to subtract.
These deltas are subtraction, not new judging. The v1.7.0 report deliberately stopped short of case-level deltas. This page takes them by subtracting the two arms of each already-judged pair. No output was re-judged and no model was refitted.
Cases are not interchangeable. Each mode is tested on the jobs it exists for: write on drafting and reworking, audit on finding defects, digest on reporting what was read, viz on showing an existing structure. A single pooled score would average these into a number that describes none of them.
Read the full methodology and limitations Back to the Minto home page