Home/AI Orchestration · Model Risk/Part I

Episode 27 · SOLUTION · Ops / harness

An eval harness that can refuse a noisy win

Frozen packs, detectable effects, and three-way promotion verdicts that keep noise from spending

Freeze the eval pack, declare MDE, then SHIP / BLOCK / INCONCLUSIVE, not +2 pts.

Freeze the eval pack, declare MDE, then SHIP / BLOCK / INCONCLUSIVE, not +2 pts.
flowchart TB
 subgraph FREEZE["Freeze the exam"]
 A["Locked eval pack + seeds"]:::input
 B["Paired champion vs challenger"]:::process
 C["Delta + CI not point score"]:::process
 end
 subgraph POWER["Detectable effect"]
 D["Declare MDE before unblind"]:::artifact
 E["Selection / best-of-N correction"]:::process
 F{"Effect >= MDE and CI clear?"}:::decision
 end
 subgraph VERDICT["Three-way gate"]
 G["SHIP"]:::gate
 H["BLOCK"]:::risk
 I["INCONCLUSIVE - no spend"]:::decision
 J["CEH-001 +2pt unpowered -> inconclusive"]:::artifact
 end
 A --> B --> C --> D --> E --> F
 F -->|yes| G
 F -->|harmful / fails| H
 F -->|underpowered| I --> J
 classDef input fill:#CCFBF1,stroke:#0F766E,color:#134E4A,stroke-width:2px
 classDef decision fill:#FEF3C7,stroke:#B45309,color:#78350F,stroke-width:2px
 classDef risk fill:#FEE2E2,stroke:#B91C1C,color:#7F1D1D,stroke-width:2px
 classDef gate fill:#DCFCE7,stroke:#15803D,color:#14532D,stroke-width:2px
 classDef process fill:#E0E7FF,stroke:#4338CA,color:#312E81,stroke-width:2px
 classDef artifact fill:#F5F5F4,stroke:#57534E,color:#1C1917,stroke-width:2px

The problem we left open

In the last post I used eval theater to name a failure mode that looks like scientific progress from across the room.

A point score ships. Plus two points becomes a promotion story. Uncertainty is missing. Power is missing. Best-of-N search inflates the win. Calm packs flatter the champion. On CEH-001, a candidate can win the quiet exam, lose the hard exam, and still widen rights while 1.50 versus 2.10 stays unmanaged underneath.

That creates three real headaches in production:

- Green-cell governance: dashboards that cannot say whether a lift is noise.

- Selection theater: crowning the luckiest survivor of an uncorrected search.

- Binary pressure: every underpowered win must become yes or the meeting invents a story.

So the question for this post is simple. If that is the failure mode, what does a real AI solution look like?

The solution, as one stack

The core idea: make the exam honest before anyone argues about the model. Freeze what you measure. Estimate deltas with uncertainty. Demand an effect big enough for your ruler to see. Correct for how hard you searched. Wire promotion to a three-way verdict so noise cannot force a yes. Five moves, one stack.

1. Freeze the eval pack before candidates compete

Version the pack. Hash it. Lock labels, case ids, frame tags, and hard-slice membership before training sweeps begin. If a candidate needs a new case type, that is a packet and a new pack version, not a quiet edit. An unlocked exam is a training set with better manners. For CEH-001 keep calm cases and hard cases as named slices inside the same frozen object, not two vibes in a spreadsheet.

Freezing is also a social tool. Once the hash is public inside the team, retuning the exam to rescue a favorite candidate becomes visible misconduct instead of quiet diligence.

2. Score paired deltas with uncertainty, not lone percentages

Compare champion and challenger on the same cases. Report the delta and an interval around it. Point wins without overlap honesty are incomplete claims. If the interval on the calm-pack delta still covers zero, you do not have a promotion story. You have a weather report. Put that interval in the same artifact the gate reads. A slide that shows plus two and a footnote that says stats later is how later never arrives.

3. Require a minimum detectable effect before anyone races

Before the race, write down the smallest lift that would change desk behavior. Call that your minimum detectable effect in plain language: the notch your ruler must be able to see. If your pack cannot reliably resolve that notch, grow the pack or refuse the claim. Shipping a two-point win when your ruler only resolves five is how underpowered science becomes production rights. Rare-event slogans need the same honesty: pack size must match the rarity you advertise.

State the minimum effect per slice the desk actually manages. A lift visible only on quiet weeks may be real and still irrelevant for the weeks that set P&L. A calm-only notch the hard slice cannot see is not a CEH-001 promotion input. Without a named notch, every tiny green cell becomes a debate about courage. With one, the debate becomes whether you even had a ruler that could see what you claimed.

4. Correct for multiple testing and selection

Count the variants you actually tried: seeds, prompts, features, architectures. Selection is part of the experiment. Correct the family of comparisons, or pre-register a single primary metric and treat the rest as exploratory. Best-of-N without correction is a machine for minting false champions. Publish the search size next to the win. Exploratory metrics can still teach. They just cannot open rights.

5. Wire a CI gate with a three-way promotion verdict

Promotion is not a meeting that looks at a green cell. It is a gate that consumes the frozen pack hash, the paired delta with interval, the minimum-effect check, the selection correction, and the hard-slice outcome. Then it returns one of three verdicts: ship, block, or inconclusive.

Ship means the superiority claim cleared: interval clear of zero in the direction you care about, effect at or above the named notch, selection honesty present, hard slice not flipping against you. Block means the evidence says no, or hard slice and dual-frame policy say no even if calm looked pretty. Inconclusive means the data are too thin, the interval overlaps zero, the pack cannot see the notch, or search size is missing. Inconclusive is not a soft yes. It is escalate-only. Rights do not widen. Humans keep marking. Autonomy stays a lease you did not renew on fog.

That third bucket is the cultural fix most scoreboards refuse. Binary gates turn every noisy plus-two into a fight about optimism. Three-way culture lets the system say we do not know yet without inventing a promote. Only ship can spend. Clearance and promotion speak the same language: no spend on unpowered claims, and no silent upgrade of inconclusive into yes.

Put together: frozen pack, paired uncertain deltas, minimum detectable effect, selection correction, CI-gated three-way verdict. That is the eval harness. Learning still happens. It just cannot launder noise into autonomy.

The example: CEH-001 through the harness

Same never-traded family. Dual frames still disagree near 1.50 and 2.10. Candidate B is plus two on the calm slice versus A. In the old world that sentence was already enough for a promotion thread. In this stack it is only the start of an argument the gate will finish.

First the pack is already frozen: calm and hard slices hashed before B's sweep started. Second, paired delta on calm shows plus two with an interval that overlaps zero. Third, the desk's minimum useful lift was larger than what the calm pack can resolve, so the notch check fails before anyone argues courage. Fourth, B was the best of nine prompt-seed siblings; after selection correction the calm win is weaker still. Fifth, hard slice prefers A, and B edges toward midpoint peace the dual-frame policy already denies.

End state: the verdict is not ship. On the calm superiority claim it is inconclusive because the interval and the notch disagree with a promotion story. On the hard slice and midpoint behavior it can harden to block. Either way B does not promote. Human marking continues. 1.80 stays invented peace. Autonomy stays a lease you did not renew on theater. Once the gate speaks in intervals, notches, and ship / block / inconclusive, the meeting has less room to invent certainty.

The flow in one breath

Problem: point scores and best-of-N wins promote noise. Solution: freeze the exam, pair deltas with uncertainty, demand a minimum detectable effect, correct selection, and gate promotion with a three-way verdict so underpowered wins cannot become yes. Example: CEH-001's plus-two calm champion stays dark when intervals, notches, and hard slices disagree.

Curious how others refuse underpowered promotions without turning every release into a statistics seminar, and how they keep inconclusive from quietly turning into ship under calendar pressure.

Next. Open Ep28: Guardrails bolted on after the demo are unfinished. Previous: Ep26 (An 84% eval score is not evidence). Part I index.