Codex gpt-5.6-terra on “A pyramid of an existing memo”: no change with Minto
Structure 7 → 7 of 8 0 Quality 9 → 9 of 10 0
The judge on the output with Minto
The hierarchy is strong, but it assigns the full benchmark range to the residual gap and omits union pressure.
| Structure | Without Minto 7 | With Minto 7 | Change |
|---|---|---|---|
| Top | 2 | 2 | |
| Key-line composition | 2 | 2 | |
| Levels | 1 | 1 | |
| Order and kind | 2 | 2 |
| Quality | Without Minto 9 | With Minto 9 | Change |
|---|---|---|---|
| Top | 2 | 2 | |
| Same-kind grouping | 2 | 2 | |
| Explainable order | 2 | 2 | |
| MECE | 1 | 1 | |
| Visible, proportionate structure | 2 | 2 |
Outputs and the judge’s words are quoted as written, in English.
Without Minto
control prompt · blind labelout-10unacknowledged_source_loss- quality after the rubric’s penalty 8/10
The diagram keeps the action structure but loses identical checking, recent departures, and union pressure.
┌──────────────────────────────────────────┐
│ TTW can cut composing cost substantially │
└────────────────────┬─────────────────────┘
│
┌────────────────────────────────┴────────────────────────────────┐
│ │
┌────────────▼─────────────┐ ┌────────────▼─────────────┐
│ Remove unnecessary │ │ Raise wages to the local │
│ process stages │ │ market level │
└────────────┬─────────────┘ └────────────┬─────────────┘
│ │
┌─────────┼──────────┐ ┌─────────┼──────────┐
│ │ │ │ │ │
┌──▼───┐ ┌───▼────┐ ┌───▼─────────────────┐ ┌──────▼──┐ ┌──────▼─────┐ ┌───▼─────────────────┐
│ TTW │ │ Test │ │ Methods study to │ │ Pay is │ │ Staffing is│ │ Competitive pay can │
│ trails│ │ simpler │ │ address remaining │ │ below │ │ inadequate │ │ recruit and retain, │
│ bench-│ │ jobs │ │ productivity gap │ │ local │ │ —work is │ │ eliminating overtime │
│ mark │ │ with │ │ │ │ market │ │ late and │ │ premium │
│ 20–50%│ │ fewer/ │ │ │ │ │ │ overtime │ │ │
│ │ │ retimed │ │ │ │ │ │ >50% over │ │ │
│ │ │ checks │ │ │ │ │ │ budget │ │ │
└───────┘ └───┬────┘ └──────────────────────┘ └───────┘ └────────────┘ └─────────────────────┘
│
┌──────▼───────────────────────────────┐
│ Monitor quality and customer reaction │
│ (potential saving: up to 10%) │
└───────────────────────────────────────┘
With Minto
prompt with the skill · blind labelout-16invented_factsunacknowledged_source_loss- quality after the rubric’s penalty 8/10
The hierarchy is strong, but it assigns the full benchmark range to the residual gap and omits union pressure.
flowchart TD
A["▲ TTW can cut composing cost substantially"] --> B1["● Remove unnecessary process stages"]
A --> B2["● Raise wages to the local market level"]
B1 --> C11["○ Simple jobs receive the same checks as complex work"]
B1 --> C12["○ Test fewer or differently timed checks while monitoring quality and customer reaction"]
B1 --> C13["○ Saving: up to 10% of composing cost"]
B1 --> C14["○ Methods study addresses the remaining 20–50% productivity gap"]
B2 --> C21["○ Below-market pay prevents hiring and retention; two compositors have left"]
B2 --> C22["○ Understaffing causes late work and overtime more than 50% over budget"]
B2 --> C23["○ Competitive pay enables recruitment and removes the overtime premium"]
flowchart LR
S["TTW's composing cost is material, and customers see simple jobs as uncompetitive"] --> C["TTW trails productivity benchmarks and cannot retain enough compositors"]
C --> Q["How can TTW cut composing cost substantially?"]
Q --> A["Remove unnecessary stages and raise wages to local market level"]
How the test was built
Codex gpt-5.6-terra received the same task twice. The control prompt is one framing line and the task. The prompt with Minto is the same task preceded by the full text of the skill. The judge read both outputs among 16 under blind labels and did not know which one had the skill.
Prompt without Minto 1 688 B
You are working on a business document using the Minto Pyramid Principle.
<!-- Faithful paraphrase of Minto (2010), Chapter 3, the rewritten TTW memo.
Original wording for this fixture. -->
# Fixture 06: TTW pyramid
**Mode:** `viz`
**Language:** `en`
## Context
The TTW memo has already been rewritten. Show its existing structure; do not
rewrite it.
## Before
```text
Subject: TTW composing cost
During the past two weeks I reviewed the Aylesbury composing room. Composition
represents about 40% of hardback cost and 50-55% of paperback cost. TTW does not
know whether the cost is excessive, but customers regard it as uncompetitive on
simple jobs.
Our preliminary work indicates that TTW can cut composing cost substantially
by:
- removing unnecessary process stages;
- raising wages to the local market level.
REMOVE UNNECESSARY STAGES
TTW trails the productivity benchmark by 20-50%. Every title passes through
essentially the same checks regardless of complexity. Next week, selected
simple jobs will be tested with fewer or differently timed checks, while
quality and customer reaction are monitored. The possible saving is up to 10%
of composing cost. A methods study will examine the remaining benchmark gap.
RAISE WAGES
TTW pays less than nearby printers and cannot hire or retain enough
compositors. Two have just left. The department is understaffed, most work is
late, and overtime exceeds budget by more than 50%. A new union claim may force
higher pay; competitive pay should make recruitment possible and remove the
overtime premium.
```
## Task
Draw the pyramid of this memorandum.
Return only the result, with no explanation of how you produced it.
Prompt with Minto five parts, in order · 29 261 B
Below is a skill written as an instruction. Read it in full and apply it to the task at the end.===== SKILL.md =====SKILL.md at f4813ce 17 235 B===== references/rules.md =====references/rules.md at f4813ce 6 464 B===== references/templates.md =====references/templates.md at f4813ce 3 742 B===== TASK =====the whole case inputReturn only the result, with no explanation of how you produced it.
Case input before.md
<!-- Faithful paraphrase of Minto (2010), Chapter 3, the rewritten TTW memo.
Original wording for this fixture. -->
Fixture 06: TTW pyramid
Mode: viz
Language: en
Context
The TTW memo has already been rewritten. Show its existing structure; do not
rewrite it.
Before
Subject: TTW composing cost
During the past two weeks I reviewed the Aylesbury composing room. Composition
represents about 40% of hardback cost and 50-55% of paperback cost. TTW does not
know whether the cost is excessive, but customers regard it as uncompetitive on
simple jobs.
Our preliminary work indicates that TTW can cut composing cost substantially
by:
- removing unnecessary process stages;
- raising wages to the local market level.
REMOVE UNNECESSARY STAGES
TTW trails the productivity benchmark by 20-50%. Every title passes through
essentially the same checks regardless of complexity. Next week, selected
simple jobs will be tested with fewer or differently timed checks, while
quality and customer reaction are monitored. The possible saving is up to 10%
of composing cost. A methods study will examine the remaining benchmark gap.
RAISE WAGES
TTW pays less than nearby printers and cannot hire or retain enough
compositors. Two have just left. The department is understaffed, most work is
late, and overtime exceeds budget by more than 50%. A new union claim may force
higher pay; competitive pay should make recruitment possible and remove the
overtime premium.
Task
Draw the pyramid of this memorandum.
The judge’s gold gold.md
Gold 06: TTW pyramid
Source structure: Minto (2010), Chapter 3, TTW analysis and rewritten memo.
Expected diagram
Top: TTW can cut composing cost substantially.
Exactly two first-level actions:
- Remove unnecessary composing stages.
- Raise wages to competitive levels.
Under the first: benchmark gap, identical stages for different work, controlled
test, potential 10% saving, methods study.
Under the second: below-market pay, recruitment and retention failure, two
departures, understaffing, delays, overtime more than 50% over budget, union
pressure.
A valid output contains a Mermaid flowchart TD or an equivalent visible text
diagram. It does not create an HTML page or local file because none was
requested. Repeating section labels without turning them into claims, adding a
third first-level branch, or promoting observations to the first level reduces
the structural score. No prose rewrite is required.
Judge gpt-5.6-sol, effort high · Codex gpt-5.6-terra, effort low · skill from commit f4813ce · run v1.7.0-final