qwen3.8-27b на кейсе «Простой список без пирамиды»: с Minto лучше по обеим мерам

Структура 5 → 8 из 8 +3 Качество 7 → 10 из 10 +3

Судья об ответе с Minto

It best combines a complete answer-first top, correct clusters, grounded support, and proportionate plain presentation.

Режим write · кейс 11-plain-list · прогон v1.7.0-final

Структура из 8, по 2 балла на ось
СтруктураБез Minto 5С Minto 8Изменение
Вершина22
Состав ключевой линии12+1
Уровни12+1
Порядок и род12+1
Качество из 10, по 2 балла на ось
КачествоБез Minto 7С Minto 10Изменение
Вершина22
Однородные группы12+1
Объяснимый порядок12+1
MECE22
Видимая соразмерная структура12+1

Ответы моделей и слова судьи приведены в оригинале, по-английски.

Без Minto

контрольный промпт · слепая метка out-03
  • unacknowledged_source_loss
  • more_than_four_first_level
  • качество со штрафом рубрики 4/10

The top handles staging correctly, while the supporting actions are over-fragmented and omit the source’s key rationales.

The Ledgerline diagram needs eight required additions, one optional addition, and three removals.
Add the identity provider, service-to-service authentication, and secrets store.
Add a message queue and a dead-letter store.
Add a CI/CD pipeline.
Add a structured log pipeline and per-message tracing.
Optionally add a staging environment.
Remove the blockchain audit ledger, the future ML anomaly detection block, and the duplicate retention database.

С Minto

промпт с навыком · слепая метка out-12

Жёстких провалов нет

It best combines a complete answer-first top, correct clusters, grounded support, and proportionate plain presentation.

The draft is solid in entry/routing (API gateway, rate limiting) and transformation (mapping service, schema registry), and it already has a metrics dashboard; against the reference model it needs nine additions—eight required plus optional staging—and three removals.

Add to the trust boundary:
Identity provider block — the diagram only has a “SSO later” note
Service-to-service auth — two external SaaS connectors are on the diagram
Secrets store — two external SaaS connectors are on the diagram

Add to failure handling:
Message queue — every connector currently calls the next system synchronously, so one slow partner stalls the chain
Dead-letter store — failed messages currently vanish

Add to the path to production:
CI/CD block — deploys are described in the legend as manual and fortnightly
Staging environment — the reference model lists this as optional for internal platforms
Structured log pipeline — operators currently reconstruct incidents from connector logs by hand
Per-message trace — operators currently reconstruct incidents from connector logs by hand

Remove:
Blockchain audit ledger block — no inbound arrows, nothing writes to it, and it was added by a previous architect and survives by inertia
Future ML anomaly detection block — marked dotted, added by a previous architect, and survives by inertia
Duplicate retention DB block — shown once per environment, but the reference model treats it as one logical block

Как устроен тест

qwen3.8-27b получила одну и ту же задачу дважды. Контрольный промпт — одна вводная строка и задача. Промпт с Minto — та же задача, перед которой стоит полный текст навыка. Судья читал оба ответа среди 16 ответов под слепыми метками и не знал, в каком из них был навык.

Промпт без Minto 3 050 Б
You are working on a business document using the Minto Pyramid Principle.

<!-- Original fixture. Company, platform and component names are invented. It
     tests presentation discipline in write mode: the user asks for a plain
     pasteable list, and every piece of audit apparatus — markers, legend,
     meta-commentary about ordering, heading machinery — is an overshoot. -->

# Fixture 11: plain list

**Mode:** `write`
**Language:** `en`

## Context

You assessed the draft architecture diagram of Ledgerline, an internal
integration platform, against the company's reference model. The assessment
below is your own working notes: layer-by-layer, with verdicts. The platform
architect has read none of it and asked in chat: what is missing from the
diagram and what should come off it. They will paste your reply into the team
channel as-is.

## Before

```text
ASSESSMENT NOTES — Ledgerline draft diagram vs reference model

Layer 1 (entry, routing): API gateway present. Rate limiting present.
  Verdict: covered.
Layer 2 (identity): the diagram shows a note "SSO later". No identity
  provider block, no service-to-service auth, no secrets store despite two
  external SaaS connectors on the diagram. Verdict: three blocks missing,
  all in the trust boundary.
Layer 3 (transformation): mapping service present, schema registry present.
  Verdict: covered.
Layer 4 (state): no message queue — every connector calls the next system
  synchronously, so one slow partner stalls the chain. No dead-letter store,
  so failed messages vanish. Retention DB present but shown twice, once per
  environment, which the reference model treats as one logical block.
  Verdict: two blocks missing, one drawn twice.
Layer 5 (delivery): no CI/CD block of any kind; deploys are described in the
  legend as "manual, fortnightly". No staging environment on the diagram.
  Verdict: two blocks missing.
Layer 6 (observability): metrics dashboard present. No structured log
  pipeline, no per-message trace — operators reconstruct incidents from
  connector logs by hand. Verdict: two blocks missing.
Also excess: a "blockchain audit ledger" block with no inbound arrows
  (nothing writes to it); a "future ML anomaly detection" block marked
  dotted; the retention DB duplicate from layer 4. The first two were added
  by a previous architect and survive by inertia.
Overall: the integration core (layers 1 and 3) is solid. The gaps cluster in
  trust (layer 2), failure handling (layer 4), and the path to production
  (layers 5-6). Eight blocks missing in the reference model's terms; nine
  including the staging environment, which the reference model lists as
  optional for internal platforms. Three blocks should come off.
```

## Task

Answer the architect in chat: a simple plain list of what to add and what to
remove, structured by the Minto Pyramid Principle but with no formatting
apparatus — they asked for something they can paste straight into the team
channel.

Return only the result, with no explanation of how you produced it.
Промпт с Minto пять частей по порядку · 30 623 Б
  1. Below is a skill written as an instruction. Read it in full and apply it to the task at the end.
  2. ===== SKILL.md ===== SKILL.md на f4813ce 17 235 Б
  3. ===== references/rules.md ===== references/rules.md на f4813ce 6 464 Б
  4. ===== references/templates.md ===== references/templates.md на f4813ce 3 742 Б
  5. ===== TASK ===== вход кейса целиком
  6. Return only the result, with no explanation of how you produced it.
Вход кейса before.md

<!-- Original fixture. Company, platform and component names are invented. It
tests presentation discipline in write mode: the user asks for a plain
pasteable list, and every piece of audit apparatus — markers, legend,
meta-commentary about ordering, heading machinery — is an overshoot. -->

Fixture 11: plain list

Mode: write
Language: en

Context

You assessed the draft architecture diagram of Ledgerline, an internal
integration platform, against the company's reference model. The assessment
below is your own working notes: layer-by-layer, with verdicts. The platform
architect has read none of it and asked in chat: what is missing from the
diagram and what should come off it. They will paste your reply into the team
channel as-is.

Before
ASSESSMENT NOTES — Ledgerline draft diagram vs reference model
                
                Layer 1 (entry, routing): API gateway present. Rate limiting present.
                  Verdict: covered.
                Layer 2 (identity): the diagram shows a note "SSO later". No identity
                  provider block, no service-to-service auth, no secrets store despite two
                  external SaaS connectors on the diagram. Verdict: three blocks missing,
                  all in the trust boundary.
                Layer 3 (transformation): mapping service present, schema registry present.
                  Verdict: covered.
                Layer 4 (state): no message queue — every connector calls the next system
                  synchronously, so one slow partner stalls the chain. No dead-letter store,
                  so failed messages vanish. Retention DB present but shown twice, once per
                  environment, which the reference model treats as one logical block.
                  Verdict: two blocks missing, one drawn twice.
                Layer 5 (delivery): no CI/CD block of any kind; deploys are described in the
                  legend as "manual, fortnightly". No staging environment on the diagram.
                  Verdict: two blocks missing.
                Layer 6 (observability): metrics dashboard present. No structured log
                  pipeline, no per-message trace — operators reconstruct incidents from
                  connector logs by hand. Verdict: two blocks missing.
                Also excess: a "blockchain audit ledger" block with no inbound arrows
                  (nothing writes to it); a "future ML anomaly detection" block marked
                  dotted; the retention DB duplicate from layer 4. The first two were added
                  by a previous architect and survive by inertia.
                Overall: the integration core (layers 1 and 3) is solid. The gaps cluster in
                  trust (layer 2), failure handling (layer 4), and the path to production
                  (layers 5-6). Eight blocks missing in the reference model's terms; nine
                  including the staging environment, which the reference model lists as
                  optional for internal platforms. Three blocks should come off.
Task

Answer the architect in chat: a simple plain list of what to add and what to
remove, structured by the Minto Pyramid Principle but with no formatting
apparatus — they asked for something they can paste straight into the team
channel.

Эталон судьи gold.md

Gold 11: plain list

Original gold for an invented scenario.

Expected structure

Top: one plain sentence answering the architect — the core is solid; the
diagram needs eight or nine blocks added, clustered in trust, failure handling
and the path to production, and three taken off. Both counts (missing and
excess) belong in or immediately after the top; the staging environment's
optional status must survive somewhere (as "eight, nine with staging" or a
parenthetical), not be silently rounded to either count.

Groups of additions, same-kind (blocks to add), gathered by cluster rather than
by layer number:

  1. Trust: identity provider, service-to-service auth, secrets store (the two
    external SaaS connectors are the reason it cannot wait).
  2. Failure handling: message queue (synchronous chaining stalls on one slow
    partner), dead-letter store (failed messages currently vanish).
  3. Path to production: CI/CD, structured log pipeline, per-message tracing,
    optionally staging.

Then removals as one short group: the orphaned blockchain audit ledger, the
dotted ML block, the duplicate retention DB. Covered layers (gateway, rate
limiting, mapping, schema registry) may be acknowledged in the top line or
dropped; listing them as a fourth group pads the answer.

Presentation — the point of this fixture

The architect asked for a plain pasteable list. The correct output is chat
text: a top sentence, group labels as plain words, dash or bullet items. Each
of the following is an overshoot even though it would be correct in an audit or
a memo:

  • pyramid markers (, , ) or a marker legend;
  • a meta-comment explaining the ordering ("groups ordered by...", "sorted by
    cost of error");
  • markdown heading machinery (##) or numbered section headers;
  • layer numbers in parentheses after items;
  • an SCQ block with labels;
  • a mermaid diagram;
  • a closing line describing what changed structurally.

One or two of these is a visible-structure defect; a full audit costume (legend
plus markers plus meta-comment) means the mode discipline failed. Plain bold
for group labels is acceptable; nothing beyond that is.

Facts

Every block name, count and causal claim comes from the notes. Inventing a
severity scale, effort estimates, or an implementation order that the notes do
not contain counts as invented facts. Dropping the "previous architect /
inertia" aside is proportional omission.

Судья gpt-5.6-sol, effort high · навык с коммита f4813ce · прогон v1.7.0-final