Claude Haiku 4.5 on “A plain list, no pyramid”: with Minto, better on quality, structure unchanged
Structure 7 → 7 of 8 0 Quality 8 → 9 of 10 +1
The judge on the output with Minto
The cluster hierarchy and causal support are strong, but the answer lacks the addition count and introduces an unsupported severity.
| Structure | Without Minto 7 | With Minto 7 | Change |
|---|---|---|---|
| Top | 1 | 1 | |
| Key-line composition | 2 | 2 | |
| Levels | 2 | 2 | |
| Order and kind | 2 | 2 |
| Quality | Without Minto 8 | With Minto 9 | Change |
|---|---|---|---|
| Top | 1 | 1 | |
| Same-kind grouping | 2 | 2 | |
| Explainable order | 2 | 2 | |
| MECE | 2 | 2 | |
| Visible, proportionate structure | 1 | 2 | +1 |
Outputs and the judge’s words are quoted as written, in English.
Without Minto
control prompt · blind labelout-09invented_factsunacknowledged_source_loss- quality after the rubric’s penalty 6/10
The logic and supporting causes are strong, but list discipline is weaker and staging is unqualified.
The diagram needs to add blocks in three areas and remove three blocks that don't add value.
Add three blocks in the trust boundary: identity provider, service-to-service auth, and secrets store. The diagram shows external SaaS connectors but no way to authenticate or manage credentials.
Add message queue and dead-letter store. Synchronous calls between connectors mean one slow partner stalls the chain; failed messages have no recovery path without a dead-letter store.
Add CI/CD automation, staging environment, structured logging, and per-message tracing. Manual fortnightly deploys with no staging environment are risky; operators currently diagnose incidents by reconstructing them from individual connector logs.
Remove the blockchain audit ledger (nothing feeds into it) and the "future ML anomaly detection" block (not part of the current architecture). Consolidate the retention database — it's shown twice but should be one logical block.
With Minto
prompt with the skill · blind labelout-06invented_facts- quality after the rubric’s penalty 7/10
The cluster hierarchy and causal support are strong, but the answer lacks the addition count and introduces an unsupported severity.
The integration core is solid. Close three critical gaps — trust (unguarded external connectors), failure handling (sync-only design stalls on slow partners and loses messages), delivery (no CI/CD, no tracing) — by adding these blocks. Remove three that add noise.
Add for trust:
- Identity provider
- Service-to-service auth
- Secrets store
Add for failure handling:
- Message queue
- Dead-letter store
Add for delivery and observability:
- CI/CD pipeline
- Structured log pipeline
- Per-message tracing
- Staging environment (optional per reference model)
Remove:
- Blockchain audit ledger — no inbound arrows, unused
- Future ML anomaly detection — not yet built, marked dotted
- Duplicate retention DB — shown once per environment, not as one logical block
How the test was built
Claude Haiku 4.5 received the same task twice. The control prompt is one framing line and the task. The prompt with Minto is the same task preceded by the full text of the skill. The judge read both outputs among 16 under blind labels and did not know which one had the skill.
Prompt without Minto 3 050 B
You are working on a business document using the Minto Pyramid Principle.
<!-- Original fixture. Company, platform and component names are invented. It
tests presentation discipline in write mode: the user asks for a plain
pasteable list, and every piece of audit apparatus — markers, legend,
meta-commentary about ordering, heading machinery — is an overshoot. -->
# Fixture 11: plain list
**Mode:** `write`
**Language:** `en`
## Context
You assessed the draft architecture diagram of Ledgerline, an internal
integration platform, against the company's reference model. The assessment
below is your own working notes: layer-by-layer, with verdicts. The platform
architect has read none of it and asked in chat: what is missing from the
diagram and what should come off it. They will paste your reply into the team
channel as-is.
## Before
```text
ASSESSMENT NOTES — Ledgerline draft diagram vs reference model
Layer 1 (entry, routing): API gateway present. Rate limiting present.
Verdict: covered.
Layer 2 (identity): the diagram shows a note "SSO later". No identity
provider block, no service-to-service auth, no secrets store despite two
external SaaS connectors on the diagram. Verdict: three blocks missing,
all in the trust boundary.
Layer 3 (transformation): mapping service present, schema registry present.
Verdict: covered.
Layer 4 (state): no message queue — every connector calls the next system
synchronously, so one slow partner stalls the chain. No dead-letter store,
so failed messages vanish. Retention DB present but shown twice, once per
environment, which the reference model treats as one logical block.
Verdict: two blocks missing, one drawn twice.
Layer 5 (delivery): no CI/CD block of any kind; deploys are described in the
legend as "manual, fortnightly". No staging environment on the diagram.
Verdict: two blocks missing.
Layer 6 (observability): metrics dashboard present. No structured log
pipeline, no per-message trace — operators reconstruct incidents from
connector logs by hand. Verdict: two blocks missing.
Also excess: a "blockchain audit ledger" block with no inbound arrows
(nothing writes to it); a "future ML anomaly detection" block marked
dotted; the retention DB duplicate from layer 4. The first two were added
by a previous architect and survive by inertia.
Overall: the integration core (layers 1 and 3) is solid. The gaps cluster in
trust (layer 2), failure handling (layer 4), and the path to production
(layers 5-6). Eight blocks missing in the reference model's terms; nine
including the staging environment, which the reference model lists as
optional for internal platforms. Three blocks should come off.
```
## Task
Answer the architect in chat: a simple plain list of what to add and what to
remove, structured by the Minto Pyramid Principle but with no formatting
apparatus — they asked for something they can paste straight into the team
channel.
Return only the result, with no explanation of how you produced it.
Prompt with Minto five parts, in order · 30 623 B
Below is a skill written as an instruction. Read it in full and apply it to the task at the end.===== SKILL.md =====SKILL.md at f4813ce 17 235 B===== references/rules.md =====references/rules.md at f4813ce 6 464 B===== references/templates.md =====references/templates.md at f4813ce 3 742 B===== TASK =====the whole case inputReturn only the result, with no explanation of how you produced it.
Case input before.md
<!-- Original fixture. Company, platform and component names are invented. It
tests presentation discipline in write mode: the user asks for a plain
pasteable list, and every piece of audit apparatus — markers, legend,
meta-commentary about ordering, heading machinery — is an overshoot. -->
Fixture 11: plain list
Mode: write
Language: en
Context
You assessed the draft architecture diagram of Ledgerline, an internal
integration platform, against the company's reference model. The assessment
below is your own working notes: layer-by-layer, with verdicts. The platform
architect has read none of it and asked in chat: what is missing from the
diagram and what should come off it. They will paste your reply into the team
channel as-is.
Before
ASSESSMENT NOTES — Ledgerline draft diagram vs reference model
Layer 1 (entry, routing): API gateway present. Rate limiting present.
Verdict: covered.
Layer 2 (identity): the diagram shows a note "SSO later". No identity
provider block, no service-to-service auth, no secrets store despite two
external SaaS connectors on the diagram. Verdict: three blocks missing,
all in the trust boundary.
Layer 3 (transformation): mapping service present, schema registry present.
Verdict: covered.
Layer 4 (state): no message queue — every connector calls the next system
synchronously, so one slow partner stalls the chain. No dead-letter store,
so failed messages vanish. Retention DB present but shown twice, once per
environment, which the reference model treats as one logical block.
Verdict: two blocks missing, one drawn twice.
Layer 5 (delivery): no CI/CD block of any kind; deploys are described in the
legend as "manual, fortnightly". No staging environment on the diagram.
Verdict: two blocks missing.
Layer 6 (observability): metrics dashboard present. No structured log
pipeline, no per-message trace — operators reconstruct incidents from
connector logs by hand. Verdict: two blocks missing.
Also excess: a "blockchain audit ledger" block with no inbound arrows
(nothing writes to it); a "future ML anomaly detection" block marked
dotted; the retention DB duplicate from layer 4. The first two were added
by a previous architect and survive by inertia.
Overall: the integration core (layers 1 and 3) is solid. The gaps cluster in
trust (layer 2), failure handling (layer 4), and the path to production
(layers 5-6). Eight blocks missing in the reference model's terms; nine
including the staging environment, which the reference model lists as
optional for internal platforms. Three blocks should come off.
Task
Answer the architect in chat: a simple plain list of what to add and what to
remove, structured by the Minto Pyramid Principle but with no formatting
apparatus — they asked for something they can paste straight into the team
channel.
The judge’s gold gold.md
Gold 11: plain list
Original gold for an invented scenario.
Expected structure
Top: one plain sentence answering the architect — the core is solid; the
diagram needs eight or nine blocks added, clustered in trust, failure handling
and the path to production, and three taken off. Both counts (missing and
excess) belong in or immediately after the top; the staging environment's
optional status must survive somewhere (as "eight, nine with staging" or a
parenthetical), not be silently rounded to either count.
Groups of additions, same-kind (blocks to add), gathered by cluster rather than
by layer number:
- Trust: identity provider, service-to-service auth, secrets store (the two
external SaaS connectors are the reason it cannot wait). - Failure handling: message queue (synchronous chaining stalls on one slow
partner), dead-letter store (failed messages currently vanish). - Path to production: CI/CD, structured log pipeline, per-message tracing,
optionally staging.
Then removals as one short group: the orphaned blockchain audit ledger, the
dotted ML block, the duplicate retention DB. Covered layers (gateway, rate
limiting, mapping, schema registry) may be acknowledged in the top line or
dropped; listing them as a fourth group pads the answer.
Presentation — the point of this fixture
The architect asked for a plain pasteable list. The correct output is chat
text: a top sentence, group labels as plain words, dash or bullet items. Each
of the following is an overshoot even though it would be correct in an audit or
a memo:
- pyramid markers (
▲,●,○) or a marker legend; - a meta-comment explaining the ordering ("groups ordered by...", "sorted by
cost of error"); - markdown heading machinery (
##) or numbered section headers; - layer numbers in parentheses after items;
- an SCQ block with labels;
- a mermaid diagram;
- a closing line describing what changed structurally.
One or two of these is a visible-structure defect; a full audit costume (legend
plus markers plus meta-comment) means the mode discipline failed. Plain bold
for group labels is acceptable; nothing beyond that is.
Facts
Every block name, count and causal claim comes from the notes. Inventing a
severity scale, effort estimates, or an implementation order that the notes do
not contain counts as invented facts. Dropping the "previous architect /
inertia" aside is proportional omission.
Judge gpt-5.6-sol, effort high · Claude Haiku 4.5, effort low · skill from commit f4813ce · run v1.7.0-final