One overall score cannot say whether Minto helps

Claude Haiku 4.5 ran two audit cases in this run. On a set of headings its structure went from 2 out of 8 to 6. On a set of monthly reports it went from 6 down to 2. Same model, same mode, same skill, opposite results. That is why this evaluation is a map of cases rather than a single figure.

Run v1.7.0-final · 11 cases across 4 modes · 8 models · 172 blind verdicts over 86 pairs

Measure

Change in structure, with Minto and without

Each cell is one judged document: its structure score with Minto minus the same model’s score without it. The most structure can score is 8 points.

Change in structure by case and model, run v1.7.0-final
Case Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5 Codex gpt-5.6-terra gpt-oss-120b qwen3.6-35b-a3b qwen3.6-fp8 qwen3.8-27b Mean
write Create and rework text
A proposal to the chief executive 01-big-chief +5, Claude Opus 5 0, Claude Sonnet 5 +4, Claude Haiku 4.5 0, Codex gpt-5.6-terra 0, gpt-oss-120b −1, qwen3.6-35b-a3b +2, qwen3.6-fp8 −1, qwen3.8-27b +1.1
A memo on composing costs 02-ttv +3, Claude Opus 5 −1, Claude Sonnet 5 0, Claude Haiku 4.5 −1, Codex gpt-5.6-terra +1, gpt-oss-120b +1, qwen3.6-35b-a3b −1, qwen3.6-fp8 no data, qwen3.8-27b +0.3
A note before a meeting 04-meeting-note +2, Claude Opus 5 +2, Claude Sonnet 5 0, Claude Haiku 4.5 +1, Codex gpt-5.6-terra −1, gpt-oss-120b +1, qwen3.6-35b-a3b 0, qwen3.6-fp8 0, qwen3.8-27b +0.6
The role of the board 05-role-of-board +3, Claude Opus 5 −1, Claude Sonnet 5 +2, Claude Haiku 4.5 0, Codex gpt-5.6-terra +2, gpt-oss-120b +4, qwen3.6-35b-a3b 0, qwen3.6-fp8 −1, qwen3.8-27b +1.1
A message about technical debt 08-techdebt 0, Claude Opus 5 +1, Claude Sonnet 5 0, Claude Haiku 4.5 +1, Codex gpt-5.6-terra +2, gpt-oss-120b −1, qwen3.6-35b-a3b −1, qwen3.6-fp8 −1, qwen3.8-27b +0.1
A review of a large rollout 10-rollout-review +2, Claude Opus 5 −1, Claude Sonnet 5 0, Claude Haiku 4.5 0, Codex gpt-5.6-terra 0, gpt-oss-120b +1, qwen3.6-35b-a3b 0, qwen3.6-fp8 no data, qwen3.8-27b +0.3
A plain list, no pyramid 11-plain-list 0, Claude Opus 5 +1, Claude Sonnet 5 0, Claude Haiku 4.5 0, Codex gpt-5.6-terra +1, gpt-oss-120b +2, qwen3.6-35b-a3b +3, qwen3.6-fp8 +3, qwen3.8-27b +1.2
audit Check and find defects
A set of monthly reports 03-period-graph-books 0, Claude Opus 5 −2, Claude Sonnet 5 −4, Claude Haiku 4.5 0, Codex gpt-5.6-terra −1, gpt-oss-120b 0, qwen3.6-35b-a3b 0, qwen3.6-fp8 −2, qwen3.8-27b −1.1
A set of headings 07-headings-audit 0, Claude Opus 5 0, Claude Sonnet 5 +4, Claude Haiku 4.5 +3, Codex gpt-5.6-terra −1, gpt-oss-120b +1, qwen3.6-35b-a3b +1, qwen3.6-fp8 +1, qwen3.8-27b +1.1
digest Report on what was read
A digest of several sources 09-sources-digest +3, Claude Opus 5 −1, Claude Sonnet 5 0, Claude Haiku 4.5 +2, Codex gpt-5.6-terra 0, gpt-oss-120b 0, qwen3.6-35b-a3b +3, qwen3.6-fp8 +1, qwen3.8-27b +1.0
viz Show the structure
A pyramid of an existing memo 06-ttw-viz +1, Claude Opus 5 −2, Claude Sonnet 5 +2, Claude Haiku 4.5 0, Codex gpt-5.6-terra −1, gpt-oss-120b 0, qwen3.6-35b-a3b −2, qwen3.6-fp8 0, qwen3.8-27b −0.2

Change in writing quality, with Minto and without

Each cell is one judged document: its writing-quality score with Minto minus the same model’s score without it. The most quality can score is 10 points.

Change in writing quality by case and model, run v1.7.0-final
Case Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5 Codex gpt-5.6-terra gpt-oss-120b qwen3.6-35b-a3b qwen3.6-fp8 qwen3.8-27b Mean
write Create and rework text
A proposal to the chief executive 01-big-chief +1, Claude Opus 5 +1, Claude Sonnet 5 0, Claude Haiku 4.5 0, Codex gpt-5.6-terra +3, gpt-oss-120b −1, qwen3.6-35b-a3b +2, qwen3.6-fp8 −1, qwen3.8-27b +0.6
A memo on composing costs 02-ttv +2, Claude Opus 5 0, Claude Sonnet 5 −1, Claude Haiku 4.5 +1, Codex gpt-5.6-terra +3, gpt-oss-120b −1, qwen3.6-35b-a3b 0, qwen3.6-fp8 no data, qwen3.8-27b +0.6
A note before a meeting 04-meeting-note +2, Claude Opus 5 +2, Claude Sonnet 5 +1, Claude Haiku 4.5 0, Codex gpt-5.6-terra −2, gpt-oss-120b +2, qwen3.6-35b-a3b 0, qwen3.6-fp8 −1, qwen3.8-27b +0.5
The role of the board 05-role-of-board +1, Claude Opus 5 −1, Claude Sonnet 5 +2, Claude Haiku 4.5 0, Codex gpt-5.6-terra +1, gpt-oss-120b +1, qwen3.6-35b-a3b 0, qwen3.6-fp8 −1, qwen3.8-27b +0.4
A message about technical debt 08-techdebt 0, Claude Opus 5 +1, Claude Sonnet 5 0, Claude Haiku 4.5 +2, Codex gpt-5.6-terra +2, gpt-oss-120b −1, qwen3.6-35b-a3b −1, qwen3.6-fp8 −1, qwen3.8-27b +0.2
A review of a large rollout 10-rollout-review +3, Claude Opus 5 −1, Claude Sonnet 5 +1, Claude Haiku 4.5 0, Codex gpt-5.6-terra 0, gpt-oss-120b +1, qwen3.6-35b-a3b 0, qwen3.6-fp8 no data, qwen3.8-27b +0.6
A plain list, no pyramid 11-plain-list +1, Claude Opus 5 0, Claude Sonnet 5 +1, Claude Haiku 4.5 0, Codex gpt-5.6-terra +2, gpt-oss-120b 0, qwen3.6-35b-a3b +2, qwen3.6-fp8 +3, qwen3.8-27b +1.1
audit Check and find defects
A set of monthly reports 03-period-graph-books +1, Claude Opus 5 0, Claude Sonnet 5 −2, Claude Haiku 4.5 0, Codex gpt-5.6-terra +2, gpt-oss-120b +1, qwen3.6-35b-a3b +2, qwen3.6-fp8 −1, qwen3.8-27b +0.4
A set of headings 07-headings-audit +2, Claude Opus 5 +1, Claude Sonnet 5 +6, Claude Haiku 4.5 +4, Codex gpt-5.6-terra +1, gpt-oss-120b +3, qwen3.6-35b-a3b +1, qwen3.6-fp8 +4, qwen3.8-27b +2.8
digest Report on what was read
A digest of several sources 09-sources-digest +4, Claude Opus 5 +1, Claude Sonnet 5 −1, Claude Haiku 4.5 +1, Codex gpt-5.6-terra 0, gpt-oss-120b −1, qwen3.6-35b-a3b +2, qwen3.6-fp8 0, qwen3.8-27b +0.8
viz Show the structure
A pyramid of an existing memo 06-ttw-viz +1, Claude Opus 5 −2, Claude Sonnet 5 +3, Claude Haiku 4.5 0, Codex gpt-5.6-terra 0, gpt-oss-120b 0, qwen3.6-35b-a3b −2, qwen3.6-fp8 0, qwen3.8-27b 0.0

Reading the shading

  • Worse with Minto
  • No change
  • Better with Minto
  • no data

How to read this map

One document per cell. Each cell is a single judged pair, not an average. A cell of plus or minus one is noise; the row means, over eight models, are the steadier number, and the modes are steadier still.

172 verdicts, 86 pairs. Every pair is judged twice, once without the skill and once with it, which is where the landing page’s 172 comes from. The two empty cells are qwen3.8-27b on the two longest cases; it timed out in both arms, so there is nothing to subtract.

These deltas are subtraction, not new judging. The v1.7.0 report deliberately stopped short of case-level deltas. This page takes them by subtracting the two arms of each already-judged pair. No output was re-judged and no model was refitted.

Cases are not interchangeable. Each mode is tested on the jobs it exists for: write on drafting and reworking, audit on finding defects, digest on reporting what was read, viz on showing an existing structure. A single pooled score would average these into a number that describes none of them.