The problem we left open
In the last post I used endogenous risk to name the failure mode that survives even careful live breakers.
Predictions change the data you retrain on. Agents mint numeric claims without receipts. On CEH-001, a fluent 1.80 peace can become blotter truth and then training truth. If that loop stays open, every other control you built can notarize fiction with excellent manners.
This post closes that hole and, in the same breath, says what I mean by a shippable decision system. Not a syllabus dump. A set of systems takeaways earned the hard way.
I am closing here because endogenous risk is where a careful machine can still launder fiction into tomorrow's training set. If you only ship exogenous breakers, you still leave a path for the system to teach itself theater.
The solution, as one stack
Five moves for the endogenous hole, then the chain those moves have to sit inside.
1. Anti-feedback training and eval rules
Labels that came from denies, human panic midpoints, or unverified agent marks cannot enter production training unless explicitly re-labeled under an information contract. Maintain a counterfactual or delayed-label eval pack insulated from recent self-prints. Continuous training without this rule is how 1.80 becomes data.
2. Verify claims at write time
Every numeric or source claim an agent emits for CEH-001 must point at a tool receipt, document hash, or typed artifact id. Fail closed. Fluent paragraphs without receipts do not enter audit as evidence. They enter as blocked drafts.
3. Keep an independent challenger
Builder and challenger are different roles. Challenger owns poison packs, shortcut reports, measure sensitivity, and replay diffs. Challenger cannot be graded on launch date. If the only person who can stop autonomy also gets promoted for shipping autonomy, you built theater with a title.
4. Require the full artifact chain before autonomy widens
Tiny size still needs the chain. Wider size needs cleaner weeks on that chain. Missing link means escalate-only. This is the promotion runbook living as a permanent lease, not a one-time ceremony.
5. Keep kill switches armed
Monitoring objects, fingerprint drift, change packets, and automatic revoke stay live. Anti-feedback rules do not replace breakers. They stop you from retraining on the mess breakers already caught.
Put together: quarantine self-prints, verify claims, independent challenge, chain-gated autonomy, armed revoke. That closes the endogenous hole.
Walk one CEH-001 decision after this stack. An agent draft that cites no receipt dies at write time. A human panic midpoint cannot re-enter labels without re-contracting. A challenger can still fail the release on shortcut or measure grounds. Size stays escalate-only until the chain is complete and the lease is green.
What I take from the whole machine as systems work
I started with a circle and three ways to pick a chord. The systems moral never really changed. Anonymous geometry manufactures certainty. The work is to make the geometry explicit, testable, and owned before capital moves.
On a never-traded family like CEH-001, that moral shows up as a replayable chain I will not pretend is optional:
- Lock the ruler and keep frame disagreement visible. 1.50 versus 2.10 is information. 1.80 peace is often politics.
- Probe the near-optimal crowd. Flat leaderboards are not unique decisions.
- Demand stakeable ranges and coverage that can break, not softmax swagger.
- Split think from act with typed jobs and allow tokens.
- Pin as-of clocks and joint stress. Separate columns are not independence.
- Budget paths, name worlds, cost hedges.
- Freeze exams, multi-gate releases, escalate-only until theater dies.
- Card the loss. The objective is already the strategy.
- Hunt shortcuts on purpose. Worst-group pain beats average comfort.
- Contract the information room. Clean tables can still peek.
- Remove drift before calling a policy a hedge. Incomplete books get bands, not fake points.
- Monitor in time. Autonomy is a lease. Silent model updates are new candidates.
- Stop training on your own unverified outputs. Verify claims. Keep an independent challenger.
That list is long because the failure modes are real. It is not a curriculum outline. It is what I actually need on the table before I trust an agent with size on a structure that has never traded.
If any link is missing, the honest product is dull refusal. Dull refusal is how you keep 1.50 and 2.10 from collapsing into a trained peace cult wearing a demo-day smile.
The flow in one breath
Problem: feedback and ungrounded claims teach a stack to notarize fiction. Solution: anti-feedback rules, claim verify, independent challenge, chain-gated autonomy, armed revoke. Ship meaning: a CEH-001 decision is replayable from ruler to kill switch, or it does not spend.
Curious how others define done for financial AI systems without reducing done to a bigger model or a greener demo day.
If I compress the ship meaning into one operational sentence, it is this: no capital action without a replayable chain, and no continuous training on unverified self-prints. Everything else in this series was how to make that sentence true without turning it into a slogan.
For CEH-001 that sentence means dual frames can stay alive, midpoint peace can stay denied, hedges can stay measure-honest, and agents can stay quiet when receipts are missing. Autonomy is available. It is just expensive in the right way. Expensive enough that false certainty stops being the path of least resistance.
That is the close I wanted when I started with Bertrand. Not a bigger model. A machine that knows when it does not yet deserve to decide.
My own answer stays practical. Done is when dull refusal is cheaper than false certainty, and when 1.50 versus 2.10 can remain visible without someone needing a peace number to feel finished.