Home/AI Orchestration · Model Risk/Part I

Episode 26 · PROBLEM · Ops / harness

An 84% eval score is not evidence

Why unpowered claims and best-of-N wins turn promotion into noise theater

A point score is not a voltmeter, delta, uncertainty, and power must show up.

A point score is not a voltmeter, delta, uncertainty, and power must show up.
flowchart LR
 subgraph THEATER["Eval theater"]
 A["Slide: 84pct or +2 pts"]:::input
 B["Green check -> promote"]:::risk
 C["Point score treated as voltmeter"]:::risk
 end
 subgraph MISSING["What a claim owes"]
 D["Delta estimate"]:::artifact
 E["Uncertainty / CI"]:::artifact
 F["Sample power / MDE"]:::artifact
 G{"All three present?"}:::decision
 end
 subgraph BIAS["Selection bias"]
 H["Best-of-N champion"]:::risk
 I["Calm pack flatters winner"]:::risk
 J["CEH-001 unpowered win ships noise"]:::risk
 end
 A --> B --> C
 C --> G
 D --> G
 E --> G
 F --> G
 G -->|no| H --> I --> J
 classDef input fill:#CCFBF1,stroke:#0F766E,color:#134E4A,stroke-width:2px
 classDef decision fill:#FEF3C7,stroke:#B45309,color:#78350F,stroke-width:2px
 classDef risk fill:#FEE2E2,stroke:#B91C1C,color:#7F1D1D,stroke-width:2px
 classDef gate fill:#DCFCE7,stroke:#15803D,color:#14532D,stroke-width:2px
 classDef process fill:#E0E7FF,stroke:#4338CA,color:#312E81,stroke-width:2px
 classDef artifact fill:#F5F5F4,stroke:#57534E,color:#1C1917,stroke-width:2px

I keep watching teams ship a leaderboard number that looks like evidence and is not.

The slide says eighty-four percent. Or plus two points versus last week's champion. There is a green check. Someone writes promote. Capital starts trusting a point estimate the way it trusts a voltmeter. That habit is older than modern agents. It is also how noise gets a job title.

This post is about powered claims. Not vibes. Not demo polish. Whether the improvement you are about to promote could still be sampling theater under the uncertainty you actually have.

What a claim owes you that a point score does not

A claim like this model is better on CEH-001 marks is a statement about a population, not about one lucky draw. It owes you three boring objects: an estimate of the delta, an uncertainty around that delta, and a sense of whether your sample was large enough to see the effect you care about. Without those, eighty-four percent is a costume. Plus two points is a mood.

There is a concentration mindset I keep returning to when people talk about rare-event slogans. Chebyshev-style thinking is the plain version: if you want a tight claim about how far an average can wander, you need sample size honest enough for the claim. Rare failures need more evidence than common ones. Thirty quiet days without a miss do not certify a tail. A small calm pack where the champion wins by a hair does not certify a hard week. Thin packs lie politely, and unpowered promotion just baptizes the lie.

I care about this because promotion is a rights change. Tiny size, auto allow, fewer escalations. If the rights change rides on an unpowered score, you did not validate. You held a ceremony. The blotter will not care that the ceremony used a spreadsheet with three decimal places.

There is a second reason I care. Once a point score becomes the language of go or no-go, every incentive in the building points at moving that number. People will retune features, prompts, and even the pack membership until the cell turns green. Without power and freeze rules, that is not iteration. It is negotiating with the exam.

Why scoreboard culture feels like science

Numbers end meetings. Confidence intervals prolong them. Best-of-N search looks like diligence: try twelve prompts, twelve seeds, twelve feature recipes, keep the winner. Multiple comparisons hide inside that diligence. The more doors you open, the more often noise finds one. Selection without correction inflates wins. The champion is not only the best model. It is the luckiest survivor of your search.

Holdout packs that stay unlocked while candidates evolve are another quiet cheat. If the exam can be retuned against, the score stops being an exam. It becomes a training signal with better branding. Teams still say we froze eval. They mean we stopped renaming the file.

There is also effect-size blindness. A two-point lift can be real and still too small to matter for desk risk, or too small to detect reliably with the pack you own. Minimum detectable effect is not academic fussiness. It is asking whether your ruler can see the notch you claim to have cut.

How this shows up in production systems

1. Promotion by green cell. A dashboard shows champion ahead by two points. No interval. No paired view of the same cases. No statement of what lift would even be decision-relevant. The cell is green, so the gate opens.

2. Best-of-N inflation. Twelve variants, one winner, one story. Nobody publishes how many siblings lost. Nobody corrects for the search. The win rate on the slide is not the win rate on a fresh pack.

3. Calm-pack theater. The eval set is last quarter's quiet weeks. Hard co-break weeks are under-represented. The champion looks stable. Live stress flips the ranking. Monitoring later discovers what the pack never asked.

4. Rare-event slogans without teeth. We never saw a dual-frame collapse on the eval pack becomes a safety claim. The pack never contained the collapse mode. Absence of evidence wore a certificate.

A CEH-001 promotion that looks like progress

Keep the never-traded structure. Spread-frame work still lives near 1.50. Curve-frame work still lives near 2.10. The desk cares about mark quality under both frames, crowd width, and whether a candidate invents 1.80 peace.

Candidate B beats champion A by two points on a calm mean-error pack. Slide says promote. Look closer. Pair the same cases. Build an uncertainty band on the delta. The band overlaps zero. The two-point win is compatible with no real lift. On a harder pack with co-breaks and missing-curve days, the ranking flips: A is quieter and less wrong on the weeks that hurt. Dual-frame disagreement is still wide. B is slightly better at sounding settled near a midpoint culture the desk already mistrusts.

If you promote on the calm point score, you ship a candidate whose only clear talent may be winning an underpowered beauty contest. Rights expand. The 1.50 versus 2.10 object gets a new owner that was never shown to be better where it matters. That is eval theater with a blotter.

Notice what the two-point story erased. It erased whether the same cases drove the win. It erased whether nine siblings were discarded. It erased whether hard weeks were in the room. It erased whether a lift that small should change autonomy at all. Erasure is the product. The percentage is just the packaging.

What stacks quietly assume

Many stacks assume a point metric on a fixed CSV is already a scientific claim. They assume selecting the best of many runs is free. They assume quiet history represents stress. They assume a percentage is automatically powered because it has a percent sign. On thin structured products and shifting regimes, those assumptions are wishes.

Another quiet assumption: that challenger review can eyeball a two-point delta and know it is real. Humans are bad at that under delivery pressure. The harness has to make underpowered wins fail closed, or Friday will keep inventing champions.

What I take from this as systems work

I do not trust a promotion score that cannot show uncertainty, effect size, and selection honesty on the pack that matches the decision. In engineering language: freeze the exam before candidates compete, report deltas with intervals, refuse lifts below what you can detect, correct for best-of-N search, and keep hard packs that can flip a calm winner.

Building reliable AI infrastructure, at least for me, is less about chasing another green cell and more about refusing to let noise rent autonomy.

Curious where others still promote on point scores, and what broke first when a calm-pack champion met a hard week.

Next. Open Ep27: An eval harness that can refuse a noisy win. Previous: Ep25 (Done is a replayable chain, not a bigger model). Part I index.