I keep seeing hedge engines that look careful in a simulator and then reveal they were partly just betting the drift.
They named worlds. They turned costs on. They budgeted path failures. Charts look like risk management. Then you ask which probability world the policy was trained under, and the room goes quiet. That silence matters. The measure you train under decides whether the policy is hedging or hunting drift.
Black-Scholes is the famous clean version of that story: in a complete market you can change measure, get a unique price, and get a unique replicating hedge. That uniqueness is the textbook promise, not a universal property of every desk book. Live structured books around something like CEH-001 are usually incomplete. Frictions exist. Not every risk is tradable. Uniqueness dies. An AI that learns to hedge under the physical world can smuggle statistical arbitrage into a process labeled risk management.
This is the ruler problem wearing probability clothes. Anonymous measure choice is anonymous geometry on paths.
I have sat in reviews where hedge quality meant simulator PnL under physical drift, residual honesty meant the same chart with a smoother line, and uniqueness meant the optimizer converged. None of those sentences survive contact with an incomplete book. They survive contact with a demo.
Two worlds, two jobs
The physical world is where history happened: drifts, risk premia, the messy P you sample from tapes and simulators. The pricing or hedging-clean world is a related world where tradable gains behave fairly after costs in the sense finance needs for replication stories. Changing world is not a vibe. It is a change of probability weights on paths.
If you train a policy to minimize hedging error under P without removing tradable drift, the optimizer will often take free-ish bets the simulator still offers. The curves look like hedge quality. Part of the gain is I predicted the drift. When live drift differs, the hedge collapses.
Residual honesty checks catch some of this, but only if you know to look for leftover edge that is really unremoved drift. The lecture idea underneath is simple even when the machinery is not: reweight path probabilities so tradable assets behave fairly, then ask the hedge question. Deep hedging work that removes the drift is rediscovering that discipline for black-box simulators.
I keep meeting teams who think naming World F and World S finished the job. Naming worlds without naming the measure story is still anonymous geometry. The simulator seed is not a measure card.
Incomplete markets: disagreement returns inside pricing
Even after you lock frames, many pricing measures can remain compatible with no obvious arbitrage when markets are incomplete. Different measures, different prices, different optimal hedges. CEH-001's 1.50 versus 2.10 can be frame conflict. It can also be measure conflict wearing frame clothing.
Deep learning does not dissolve that. It can pick one implicit measure by picking one simulator and one objective. Shipping that choice without a measure card is anonymous geometry all over again. Loss cards named what winning meant. Measure cards name which path weights were treated as fair.
How this shows up in AI hedge stacks
1. Simulator edge farming. A friendly world still contains mild drift. The policy harvests it. Promotion metrics glow. A stressed world with drift removed or costs dominating drift looks worse. Teams keep the friendly world and call the difference realism.
2. Midpoint laundering via hedgeability. Someone claims 1.80 is fine because a P-trained policy can hedge it in the friendly sim. The hedge is partly a bet. The peace number borrows false legitimacy from a trading strategy wearing hedge clothing.
3. Greeks cosplay. A network outputs delta-like actions that match textbook intuition on average while encoding measure-specific bets in the residuals. Dashboards show Greek stability. P&L shows something else on the week drift flips.
4. Incomplete-market denial. Teams demand a unique AI price for CEH-001 after locking frames. Completeness is an assumption, not a mood. If the book is incomplete, uniqueness was never on offer. Bands are honesty. Points are theater.
Walk the failure without pretending it is exotic
Train a hedge under P with leftover drift. Strong demo PnL on the friendly world. Promote as the CEH-001 hedge. Live drift differs, or costs dominate. Residual edge flips sign. Confidence theater returns. People say the model broke. The training question was wrong.
Notice how this couples to earlier headaches without needing a syllabus. An anonymous loss can reward drift hunting. A shortcut cue can correlate with drift in the sim. An illegal information room can make the hunt look smarter. Measure confusion is another way to mint false certainty.
The painful part is social. A drift-farming policy feels like skill while it works. People defend it with Greek plots. Then drift flips and the defense becomes regime change. Sometimes it is regime change. Often it is a bet that stopped paying.
What I take from this as systems work
I do not trust hedge autonomy that cannot name its measure story. In engineering language: write a measure card, put a remove-the-drift or near-martingale stage before capital hedges, keep costs on, publish bands when uniqueness is not on offer, and bind allow scope to the measure story that was actually validated.
Another production headache sits next to incomplete-market denial: people demand a unique hedge policy because operations wants one button. Uniqueness is not an operations preference. It is a market property. If the book cannot uniquely replicate the claim, the AI cannot honestly invent uniqueness by converging. Convergence is an optimization event. Completeness is an economic claim.
So when CEH-001 shows 1.50 versus 2.10, I now ask two questions, not one. Is this frame conflict? Is this measure conflict? Either way, a midpoint blessed by a P-trained hedge is not an answer. It is a way to hide which question you refused to ask.
Curious how others separate make-money-from-the-simulator policies from hedge-the-claim policies without pretending the market is complete.