Here is a systems problem I keep running into right when a stack finally looks thick enough to ship.
You can lock rulers, probe crowds, insist on stakeable ranges, split think from act, pin as-of clocks, name joint stress, budget paths, and cost hedges. Then Friday arrives. Someone changes the exam after they see the answers. A metric leaks in. A friendly subset replaces the hard week. A joint bundle gets quietly dropped. A confidence threshold gets tuned in the war room until the candidate looks green. That move feels like delivery. It is how thin governance re-enters a thick-looking machine.
This post is not new science. It is the ship discipline that keeps the science from being negotiated away. I treat it as a promotion runbook for a never-traded family I will keep calling CEH-001: the structure where one frame keeps saying about 1.50 and another keeps saying about 2.10, and the social temptation is always the peace number in the middle.
Choosing the exam after the answers is not science
In research language this is pre-registration. On a desk it is simpler. Freeze the exam before the candidate is trained for release. Metrics, path budgets, world cards, parity packs, honesty slices, co-break bundles, and policy versions get a release id. Candidates may fail. Humans may later decide the exam was wrong. Fine. Then change the exam under a new id, and the candidate re-enters as new. What you may not do is tune the test until the model passes and call that science.
A frozen exam is a folder with a hash and an owner. If the hash changes, the release id changes. That single inconvenience blocks a remarkable amount of self-deception. Challenger versus champion is allowed. Changing the scoreboard mid-fight is not. If the challenger needs a new metric to look good, that metric joins the next frozen exam, and both candidates run again.
One green check is how theater returns
A CEH-001 family release has to clear several gates that map to the headaches I keep writing about: ruler and contract present, near-winner probe published, ranges and coverage monitor healthy, typed job schema valid, as-of parity green, joint bundles scored, trajectory budgets held on two named worlds, residual honesty clean, policy dry-run on replay. One gate fails, the build fails.
Averages across gates are forbidden. Weighted scorecards are how people waive the gate they fear. Publish the gate map next to the architecture diagram. If a gate has no owner, it will rot. If a gate can be skipped with a comment, it is not a gate. Comments belong in tickets that change the frozen exam, not in green builds. Gate owners should be people, not committees. Committees ratify. Owners get paged when a gate goes soft. Soft gates are incidents even when P&L is quiet.
Replay packs beat review-meeting eloquence
Arguments in promotion meetings are cheap. A replay pack is not: same jobs, same clocks, same worlds, same policies, new candidate versus champion. Diff the allow decisions and the act receipts. If you cannot show why allows changed, you do not promote.
Include known poison pills on purpose. A midpoint 1.80 attempt. A latest-fetch attempt that tries to sit today's exam with tomorrow's curve. A calm-only candidate that collapses under co-break. The pack should prove the gates still bite, not only that the new model looks clever on quiet days. Replay weekly even without a release. A pack that only runs under launch pressure will be under-maintained and over-negotiated.
Escalate-only until the theater dies
While CEH-001 still has dual-frame conflict, unfinished honesty, or unfinished path rulers, default posture is escalate-only for capital actions. Allow tokens require a typed job with dual frames and no midpoint, crowd width inside tolerance on calm and co-break, overlapping stakeable ranges or an explicit human frame choice, coverage breakers green, as-of parity proven for this recipe, and hedge policies scoped to two passed worlds with costs on. Missing artifact means escalate. Someone proposes 1.80 as peace means deny.
Autonomy widens only after clean replay weeks, not after a demo. That is deliberately unheroic. It will frustrate people who wanted the agent to just handle it. That frustration is the point. Unfinished geometry is not a vibe. It is a stop condition.
Staged autonomy, not a flip switch
When gates stay green, widen in stages: propose-only, then draft-plus-human-decide, then allow on tiny size, then size up inside token bounds. Each stage needs its own replay evidence. A single demo day cannot skip the ladder. Skipping the ladder is how fused think-and-act rights return wearing a launch announcement.
Canary does not mean silent. Canary means scoped tokens, loud monitoring, and a pre-agreed kill switch that does not require a meeting to fire. Meetings are how yellow dashboards become red P&L.
What I take from this as systems work
Done is not a bigger model. Done is when you can point at a CEH-001 decision and replay the chain: which ruler, which crowd, which ranges, which clock, which joints, which worlds, which allow token, which act. If that chain is missing a link, the system should sound dull and refuse. Dull refusal is the product.
I care about this because every thick stack I have seen can still be defeated by one social move: renegotiate the exam after the answers. Frozen exams, multi-gate builds, replay packs, escalate-only defaults, and staged autonomy are how you stop that move without pretending science is optional.
Curious how others enforce frozen exams across research and production without turning every release into a negotiation about which gate to waive.