The problem we left open
In the last post I named three comfort machines that show up once hedges and simulations enter the conversation.
Green averages hide the day that matters. Unnamed simulators act like another Bertrand ruler for paths. Costless hedges relocate risk into leftovers that still predict P&L. On the never-traded structure, a friendly world can make 1.80 look hedgeable while a strict world with costs and co-breaks shows the same policy leaking.
That creates headaches:
- Leaderboards that reward demo smoothness over tail honesty.
- Promotion that only survived one undeclared path measure.
- Beautiful hedge curves that assumed friction was free.
- Midpoint marks that get a second life as supposedly hedgeable peace.
So the question for this post is simple. If that is the failure mode, what does a real path-level solution look like?
The solution, as one stack
The core idea is to make the path object first-class: evaluate policies on whole trajectories against explicit failure budgets; name at least two simulation worlds and require both; optimize hedges with costs and constraints; test leftovers for leftover edge; scope Clearance allows to the worlds that passed.
I turn that into five moves. Together they are one solution, not five optional add-ons. Learning still matters. It just has to live inside named worlds and budgets instead of pretending those objects were optional polish.
1. Score the run, budget the failures
Replace mean hedge error is low with rules like: across a path pack, the number of breach days must stay under a budget, or the worst drawdown must stay under a line. Stop early when the budget is blown in offline eval. That single shift kills a lot of average-green theater because one ugly day can no longer hide in a mean.
Budgets force a conversation the mean avoids: how many bad days are we willing to automate through. If the desk cannot answer, autonomy is not ready. Silence is not a budget of zero. Silence is an unowned risk.
2. Name worlds; require two
World cards are short: dynamics assumptions, dependence bundle links, cost model, liquidity rules, seeds. Friendly-world-only success cannot promote. A second, stricter world must pass the same budgets. If the policy only lives in one ruler of paths, it does not ship.
The second world must be honestly different, not a cloned noise seed. If World S is just World F with a new random draw, you invented theater with extra paperwork. Change the dependence bundle, the cost model, or the liquidity rules enough that disagreement means something.
3. Hedge studio with costs on
Training and eval include spreads, fees, and trading limits the desk actually faces. A zero-cost delta demo can exist as a tutorial artifact. It cannot mint a Clearance token. The objective should punish churn that only looks good when friction is fake. If your hedge only works when markets are free to trade against, it is not a hedge for a desk.
4. Residual honesty checks
After the policy trades, look at leftovers. If leftovers still predict P&L in a simple way, the hedge left a free lunch on the table. Fail the pack. This is residual honesty in systems language: do not call it hedged if the trash still pays. I want that check to be as embarrassing and as mandatory as a failed unit test.
5. Scope allows to passed worlds
Clearance tokens carry which worlds and budgets passed. Live monitoring tags which world-like regime you seem to be in. If you leave the scoped set, escalate. Autonomy is conditional on the path ruler you actually validated. An allow without a world scope is another costume probability.
A concrete walkthrough
Same never-traded structure. Someone still wants a hedge near a diplomatic 1.80. Earlier rules already deny auto-posting that midpoint as a mark. Now watch the hedge path.
World F (friendly): meets mean metrics,
fails path budget on co-break week
World S (strict, costs on): breaches budget;
residual still predictive
Retrain with costs + budget objective; require both worlds
Clearance: allow hedge drafts only with token scoped to {F,S} pass
else escalate
1.80 midpoint mark still denied by earlier rules;
hedge cannot launder it
The satisfying ending is not a smoother mean curve. It is a scoped allow, or an escalate, with named worlds and budget artifacts anyone can replay. That is what a researched path-level solution means on a desk that wants to sleep.
A rollout order that avoids theater: define one failure budget the desk will actually defend, stand up World F and World S with different cost and co-break assumptions, turn costs on in the hedge objective, add one residual check that can fail the pack, then wire Clearance scope to the pass set. Do not start by training a prettier policy in an unnamed world.
When both worlds pass and leftovers look dull, you still do not auto-launder 1.80 as a mark. Earlier midpoint bans remain. The hedge stack is not a backdoor for invented peace. It is a conditional right to act inside validated path rulers.
Live regime tagging will be imperfect. That is fine. The point is not omniscience. The point is that leaving the scoped set is an escalate event, not a silent continuation. Autonomy that cannot notice it left the exam is not autonomy worth shipping.
If the desk wants one number for a slide, give them the scoped allow region and the budget status. Do not give them a lonely float and call the path problem solved.
The flow in one breath
Trajectory packs with failure budgets, then at least two named worlds, then costed hedge optimize plus residual honesty, then Clearance allow scoped to passed worlds, then live regime tag and escalate on leave. That is the systems work I care about after path theater. Curious how others keep a second world from rotting into a clone of the first, and what residual checks they trust enough to block a release.