I keep finding models that look brilliant in research and then fall over in production for a boring reason: they sat the exam with information they would not have had at decision time.
A gate can be strict and still bless garbage if the features were illegal. Crowds and ranges can look healthy if the stress pack only turns one dial. These cheats are less flashy than Bertrand. They wreck more production systems. I care about them because they survive the posts everyone likes to share: the ruler talk, the crowd talk, the swagger talk. They hide in joins and in stress catalogs.
Cheat one: the answer key arrived early
A mark at 14:00 is a claim about what the firm knew at 14:00. If a feature silently includes a 14:10 print, a revised curve published at 14:20, or a label that was only knowable at day's end, the model sat the exam with the answer key. Training metrics soar. Demo day glows. Live marking at 14:00 with only 14:00 information falls on its face, and people blame regime change instead of time travel.
Quick definition, because this gets waved away as a join bug. A capital decision has a clock. Features and tool fetches are legal only if they were knowable on that clock, not after it. As-of is the engineering word for that discipline. Agent tools make it worse. A tool call can fetch latest by default. Latest is not as-of. Latest is a hand in the cookie jar. (The deeper information-law story - what predictors may condition on - belongs in a later post. This one stays on clocks, joins, and stress catalogs.)
On a never-traded structure, the tell looks like this: the research notebook matches the afternoon revision almost perfectly. The production path, pinned to 14:00 snapshots, disagrees with both the notebook and the other frame. Teams then fix the model when they should fix the clock. I have watched this cycle burn weeks. The model was never the genius. The join was.
The painful part is how virtuous the leak feels. Someone added a better feature. Accuracy rose. A dashboard got greener. Nobody asked whether the feature was legal at decision time. In financial AI, illegal information is not a pedantic rule. It is a different exam.
Cheat two: separate shocks, independent fantasy
Desks often shock spread, then shock rates, then shock liquidity, each in isolation, then declare the book resilient. Real breaks arrive as bundles. Spreads gap while funding tightens while quotes go missing. If your AI only saw one-at-a-time pain, its crowd width and ranges are optimism machines.
We modeled them as separate is not the same claim as they are independent. Separate columns in a table do not make joint behavior vanish. Independence is a strong modeling choice. Most market distress laughs at it. When spread-frame work near 1.50 and curve-frame work near 2.10 both move together under a joint shock, the gap that looked like 0.60 in calm data can widen toward something uglier, or both clouds can jump while still refusing to overlap.
Factorized pipelines encourage the fantasy. Each service owns one driver. Each dashboard panel looks fine. Nobody owns the day when the drivers move together. That day is when the mark that spends either becomes honest or becomes an accident.
How this shows up when people trust the green dashboard
1. Notebook versus service parity theater. Offline eval used a join that production cannot reproduce at as-of time. Offline looks calibrated. Online is a different animal wearing the same model name.
2. Agent fetch discipline failure. The agent helpfully pulls the newest curve because the tool default said so. The gate sees strong ranges built on illegal information and allows. The allow was logically consistent and economically cheated.
3. Stress packs that never co-break. Hard-slice probes only perturb one factor, so you get false interchangeability across near-winners. Joint stress is where multiplicity and dependence show their teeth together.
4. False comfort from marginal intervals. Spread range looks fine. Curve range looks fine. The joint scenario that moves both is never generated, so overlap rules never see the day that matters. Green panels, empty overlap under the only stress that was real.
Walk the failure on familiar numbers
Illegal as-of
14:00 decision using 14:10 feature research mark looks accurate
pinned 14:00-only replay mark moves; swagger returns
One-at-a-time stress
shock spread alone width 0.60 feels familiar
shock curve alone similar comfort
Joint distress bundle
width between frames widens toward ~0.95
overlap still empty
auto-allow would have been a story about calm tests
The thought process I want on the desk is suspicious by default: what clock did this feature obey, and which other clocks moved with it when the world got ugly. If nobody can answer, the earlier gates are decorating a leak. More model capacity will not invent a clock. A better prompt will not invent a joint.
There is a quieter cousin of as-of leakage that shows up in labels. If you train on a mark that was only finalized at evening close, and then ask the model to emit a midday mark, you taught it to anticipate the close. The feature store can look clean while the supervision signal still time-travels.
Dependence fantasy has a dashboard cousin too. People show a correlation heatmap from calm months and declare the joint understood. Calm correlation is not co-break behavior. The bundle that matters is the one that shows up when liquidity is already thin and both frames are moving.
What I take from this as systems work
I do not trust green dashboards on thin structures until as-of discipline is forced into features and tool fetches, and until stress language can name joint bundles instead of only single knobs.
In engineering language:
- Treat as-of as a hard contract on every capital-touching feature.
- Pin agent tools to the job clock; quarantine default-latest to research-only paths.
- Fail builds when notebook and service disagree under replay.
- Re-score crowds and ranges on named co-break bundles.
- Let the gate consume co-break width, not only calm holdout smiles.
Building reliable AI infrastructure, at least for me, is less about a smarter learner and more about refusing open-book exams and one-knob courage.
Curious how others catch as-of leaks in agent tool calls before a parity incident, and whether their stress harnesses can name joint bundles rather than only single knobs.