Home/AI Orchestration · Model Risk/Part I

Episode 17 · SOLUTION · Model-build

Assume the optimizer will cheat, then make cheating expensive

Suspect cues, group packs, invariance swaps, and worst-group gates as one stack

Group packs and invariance swaps, worst-group gate, not average green.

Group packs and invariance swaps, worst-group gate, not average green.
flowchart TB
 subgraph PACKS["Group packs"]
 G1["Suspect-cue groups named"]
 G2["Regime / venue / liquidity packs"]
 BASE["Holdout is not one average"]
 end
 subgraph INV["Invariance swaps"]
 SWAP["Ablate / swap haunted features"]
 STAB["Mark must stay stable under swap"]
 FRAME["Both 1.50 and 2.10 frames tested"]
 end
 subgraph WG["Worst-group gate"]
 WORST{"Worst group within budget?"}
 PEACE{"Cue collapses to 1.80 peace?"}
 ALLOW["Promote only if worst-group holds"]
 DENY["Deny shortcut heroes"]
 end
 G1 & G2 --> BASE --> SWAP --> STAB --> FRAME
 FRAME --> WORST
 WORST -->|no| DENY
 WORST -->|yes| PEACE
 PEACE -->|yes| DENY
 PEACE -->|no| ALLOW
 classDef input fill:#CCFBF1,stroke:#0F766E,color:#134E4A,stroke-width:2px
 classDef decision fill:#FEF3C7,stroke:#B45309,color:#78350F,stroke-width:2px
 classDef risk fill:#FEE2E2,stroke:#B91C1C,color:#7F1D1D,stroke-width:2px
 classDef gate fill:#DCFCE7,stroke:#15803D,color:#14532D,stroke-width:2px
 classDef process fill:#E0E7FF,stroke:#4338CA,color:#312E81,stroke-width:2px
 classDef artifact fill:#F5F5F4,stroke:#57534E,color:#1C1917,stroke-width:2px
 class G1,G2,BASE input
 class SWAP,STAB,FRAME process
 class WORST,PEACE decision
 class ALLOW gate
 class DENY risk

The problem we left open

In the last post I used shortcut learning to name a failure mode that survives even after the objective is written down.

Models prefer easy cues that correlate with the label on the training exam. Holdouts still contain the cheat. Crowds tighten around the same Clever Hans story. On CEH-001, a quiet quote-intensity cue can pull both frames toward a fake 1.80 peace, then die in a co-break week while the real 1.50 versus 2.10 gap remains.

That creates three real headaches in production:

- The illusion of understanding: high importance on a cue that is weather, not structure.

- The arbitrage trap: agents and models agreeing because they share a haunted feature, not because the economics settled.

- Credibility collapse under shift: calm-year heroes that scramble exactly when autonomy felt earned.

So the question for this post is simple. If that is the failure mode, what does a real AI solution look like?

The solution, as one stack

The core idea is blunt: assume the optimizer will cheat if cheating is cheaper than understanding. Then make cheating expensive and detectable. I turn that into five moves that only work together.

This stack only works if the probes can fail a release. Optional notebooks do not count. If worst-group pain cannot block a token, you still have a research hobby sitting next to a production spender.

1. Maintain a suspect-cue register

List features and tool fields that are easy, regime-tied, venue-tied, or historically correlated with labels for accidental reasons. Quote intensity, venue ids, vendor composite scores, calm-only funding prints, and demo-set tool orderings belong here by default. The register is owned. It is not a wiki nobody updates.

When a new vendor field arrives looking helpful, it enters as suspect until proven otherwise. Helpful is not the same as causal. Helpful is often the cheat arriving with better packaging.

2. Build group and ablation packs before promotion

Slice evaluation by regime, venue, liquidity state, and product subfamily. Ablate suspect cues and retrain or re-infer. If performance collapses only when the cue is present, you found a shortcut, not a driver. Packs must include the co-break weeks you fear, not only random holdout rows.

I want the pack to hurt on purpose. A pack that never fails is usually a pack that never looked.

3. Run invariance swaps that should not flip the decision

Economically equivalent rewrites of inputs should leave the mark story stable: currency wrappers that should not matter, redundant encodings of the same curve information, harmless renames of venue tags when venue is not part of the contract. If a swap that should be irrelevant flips 1.50 toward 1.80, Bertrand is living inside the featurizer.

4. Gate on worst-group pain, not average comfort

Average error across quiet groups will bless shortcuts. Worst-group error, worst-regime residual, and ablated-cue deltas are the promotion numbers. A model that is beautiful on average and broken on the stressed group does not ship for capital use.

This is also where crowd probes get honest again. If every near-winner dies on the same ablated group, you do not have a tight truth cloud. You have a synchronized cheat.

5. Bind allow tokens to the shortcut report

Promotion carries a shortcut report hash: suspect cues tested, ablations run, invariance swaps passed, worst-group thresholds held. If the report is missing or stale relative to the feature recipe, no allow. That is how you stop a calm-year hero from inheriting yesterday's autonomy.

Put together: suspect register, group/ablation packs, invariance swaps, worst-group gates, report-bound allows. Learning still happens. Easy cheats just stop looking like free accuracy.

The example: killing the quote-intensity hero on CEH-001

Without the stack, quote intensity correlates with fair marks in quiet years, both frames drift toward 1.80, holdout stays green, and live co-break scrambles the book.

Now run the same case through the solution.

First the suspect register flags quote intensity and venue id as default cheats for this family.

Second, group packs score calm versus co-break. Ablation removes quote intensity. The calm group still looks fine. The co-break group was never fine without the cue. That delta is now visible before launch.

Third, invariance swaps rename venue tags and re-encode curve inputs that should be equivalent. One candidate flips toward 1.80 under a harmless rename. It fails even though average error looked sharp.

Fourth, worst-group gating refuses the calm-year hero. Dual-frame disagreement near 0.60 is preserved as managed uncertainty instead of being laundered by a haunted cue.

Fifth, allow tokens require the shortcut report. A newly added vendor composite cannot sneak in without re-running the pack. End state: you still use features. You just stop promoting weather vanes.

A practical ops note: keep the suspect register next to the feature registry so new columns cannot land in capital recipes without a probe plan. The day a helpful field arrives without an owner is the day Clever Hans gets a badge.

The flow in one breath

Problem: shortcuts pass the exam and fail the desk. Solution: register suspect cues, pack groups and ablations, swap for invariance, gate on worst groups, bind allows to the report. Example: quote-intensity peace around 1.80 dies in CI instead of on the blotter.

I also want one cultural rule around this stack. If a candidate only looks good with the suspect cue left in, that candidate is not almost ready. It is disqualified. Teams love soft language like needs more data. Sometimes that is true. Often it is a refusal to admit the model never learned the dual-frame economics and only learned summer weather.

On CEH-001 that cultural rule protects the 0.60 gap. The gap is uncomfortable. Shortcuts make it comfortable by fake peace. The stack's job is to keep the discomfort where it belongs: in managed disagreement, not in a haunted feature that collapses both frames toward 1.80 until the week quotes vanish.

Curious how others keep shortcut probes from rotting into a one-time notebook that never blocks a release.

Next. Open Ep18: A clean feature table can still sit the wrong exam. Previous: Ep16 (A model can pass the exam by learning the wrong clue). Part I index.