Home/AI Orchestration · Model Risk/Part II

Series 2 · Episode 13 · PROBLEM · Hedge · L13

Pricing with the wrong coin

Physical probabilities inside a risk-neutral formula, and the mark that comes out too rich to hedge

Physical p is the wrong coin for RN pricing.

flowchart TB
  P["Physical p"]:::input --> WRONG["Price with wrong coin"]:::risk
  Q["Risk-neutral q"]:::process --> OK["Hedge / RN price"]:::gate
  WRONG --> CODE["WRONG_MEASURE_COIN / DRIFT_OVERLAY"]:::risk

  classDef input fill:#CCFBF1,stroke:#0F766E,color:#134E4A,stroke-width:2px
  classDef decision fill:#FEF3C7,stroke:#B45309,color:#78350F,stroke-width:2px
  classDef risk fill:#FEE2E2,stroke:#B91C1C,color:#7F1D1D,stroke-width:2px
  classDef gate fill:#DCFCE7,stroke:#15803D,color:#14532D,stroke-width:2px
  classDef process fill:#E0E7FF,stroke:#4338CA,color:#312E81,stroke-width:2px
  classDef artifact fill:#F5F5F4,stroke:#57534E,color:#1C1917,stroke-width:2px

Simple claim: physical probabilities inside a risk-neutral formula mint a mark you cannot hedge. The trench is two coins in one model and the refuse when they silently swap.

I keep finding the same bug, and it never looks like a bug. It looks like someone being careful.

A pricing service takes an expectation of a discounted payoff. Somewhere upstream, a forecasting model has produced a probability that the underlying finishes up. It is a good probability, calibrated, backtested, owned. Wiring it into the pricing expectation feels like an improvement: why price with a made-up number when you have a researched one? So the service reaches for the forecast, and the mark comes out 6.7% richer than the hedge that actually discharges the claim.

Nobody notices, because the output has the right shape. It is a discounted expected payoff. It is just taken with the wrong coin.

Two probabilities that live in the same model

The one-period binomial world from last time is small enough that the confusion has nowhere to hide. S₀ = 100, up factor u = 1.1, down d = 0.9, period rate r = 0.05, and a call struck at 100 paying 10 in the up state and 0 in the down state.

There is a physical probability p. It is the real-world chance the stock goes up. It comes from history, from a view, from a model. It is what your forecaster is paid to estimate. Say p = 0.80.

There is also a risk-neutral probability q, and it does not come from anywhere near the forecaster. It falls out of the no-arbitrage structure:

q = (1 + r − d)/(u − d) = (1.05 − 0.9)/(1.1 − 0.9) = 0.75.

Notice that q is determined entirely by (u, d, r): no history, no view, no estimation. Notice also that q lies strictly between 0 and 1 exactly when d < 1+r < u, the arbitrage sandwich wearing a different hat. The condition that makes the world admissible is the condition that makes q a probability.

What q does is defined by one property. Under q, the discounted stock is a martingale:

(0.75 × 110 + 0.25 × 90)/1.05 = 105/1.05 = 100 = S₀.

Under p, it is not:

(0.80 × 110 + 0.20 × 90)/1.05 = 106/1.05 = 100.95.

That 0.95 gap is not an error. It is the risk premium, the reason anyone holds equity rather than the money market. The physical measure is supposed to have it.

So p and q are both correct, and they answer different questions. p answers what do I think will happen. q answers what does the hedge cost. The mistake is not believing in p. The mistake is putting p where q belongs.

What the wrong coin actually produces

Price the call with q, as the model intends: (0.75 × 10 + 0.25 × 0)/1.05 = 7.5/1.05 = 7.14.

Price it with p, as the "improved" service does: (0.80 × 10 + 0.20 × 0)/1.05 = 8/1.05 = 7.62.

Forty-eight cents on a 7.14 claim. Not a rounding difference, 6.7% of the premium, which on most books is several times the edge anyone claims to be capturing.

Here is what makes it more than a discrepancy. Go back to the replicating portfolio: Δ = (10 − 0)/(110 − 90) = 0.5 shares, financed by borrowing 42.86. That hedge does not contain p anywhere. It cannot, it solves two equations matching two payoff states, and probabilities appear in neither. The cost of building it is 7.14 regardless of what anyone believes about the up probability.

So 7.62 is not a slightly optimistic price. It corresponds to no portfolio at all. Anyone buying at 7.62 has handed 0.48 to whoever builds the hedge for 7.14 and delivers, with no exposure left over. The richer mark is not a stronger view expressed in price. It is a standing invitation.

And the direction is not random. Because p > q whenever the asset carries a positive risk premium, which is the normal case, the wrong coin systematically overprices calls. A book that makes this mistake does not scatter errors around the truth. It marks its long optionality consistently too high, and the mispricing grows with the very risk premium the desk is proud of forecasting.

How this shows up in production

  1. Forecast injection. A well-governed alpha model publishes a probability, and a pricing service consumes it because it is the only probability in the registry. Nobody wrote the rule that pricing may not read from that topic, because nobody imagined the two would be confused.
  1. One probability field. The instrument record has a single slot called prob_up. Two processes write to it, whichever ran last wins, and the artifact carries no field saying which measure it holds.
  1. Calibration pressure. Someone notices that q = 0.75 does not match a realized up-frequency of 0.80 and files it as a calibration defect. A ticket to "fix the pricing probability against realized outcomes," if closed as written, converts a correct model into a broken one, and the fix will look like diligence in the changelog.
  1. Blended coins. The subtlest version. Nobody swaps q for p outright; instead a drift adjustment, a sentiment tilt, or a scenario weight gets applied on top of a risk-neutral engine. The result is neither measure. It is the discrete cousin of the continuous failure: DRIFT_OVERLAY: where a physical view is smuggled into a pricing path, and it is far harder to spot than the clean substitution because no single line of code reads from the forecast.

A CEH-001 mark that looks like an improvement

Same never-traded note, same standing disagreement. The spread frame marks CEH-001 near 1.50. The curve frame marks it near 2.10. The desk denies 1.80, the midpoint is not a resolution, it is disagreement rendered invisible.

The equity-linked leg gets repriced by a service that has just been "upgraded" to use the house probability model. The leg's contribution moves up. The note's mark drifts from 1.50 toward the high 1.6s, and the memo frames this as convergence: the spread frame was stale; incorporating the forecast brings us closer to the curve view.

It is not convergence. Every basis point of that move is risk premium leaking out of a forecast and into a discounting step where it does not belong. The move is in the direction of 1.80 not because the frames are reconciling but because the overpricing bias happens to point that way, and 1.80, if it arrives, will arrive certified by a model upgrade rather than by anything the desk would recognize as a decision. The most dangerous version of the denied midpoint is not someone typing 1.80. It is a pipeline that walks there one legitimate-sounding release at a time.

Ask the only question that settles it: what does the hedge cost? The replication is untouched by the upgrade, because Δ never saw p. If the mark moved and the hedge did not, the move was not price discovery.

What stacks quietly assume

That a probability is a probability. That the best-estimated number is the right number for every expectation in the system. That an expectation operator is measure-agnostic plumbing. That a service returning a discounted expected payoff is doing risk-neutral pricing by virtue of its shape. That calibrating against realized frequencies is always an improvement. On instruments where the only defensible price is a replication cost, those assumptions are how a book gets long a bias it never chose.

What a solution must do

It has to stop treating the measure as an implementation detail and start treating it as a typed field. Every probability should be tagged with the measure it belongs to, stored as a declared (P, Q) pair rather than a single number, and refused by any consumer expecting the other one. Pricing expectations should be unable to bind a P-tagged input at all, not warned about, refused. Forecasting keeps full use of p, because that is what p is for. And any adjustment applied to a risk-neutral path should have to name itself, post its density, and explain why it is not a drift overlay.

Agents may forecast under P and may propose prices under Q. What the harness must never allow is one measure quietly doing the other's job.

Curious where others have caught a physical probability inside a pricing expectation, and whether it showed up as a P&L surprise or only as a hedge that never cost what the mark said it should.

Clearance coupling. Pricing coin must match hedge coin. Physical p inside a Q formula is WRONG_COIN refuse. Density card (later) is the receipt when coins intentionally differ.

CEH-001 equity leg marked rich under p and under-hedged under q is how wrong-coin theater becomes P&L.

Next. Open S2-14: Backward induction as the agent loop. Previous: S2-12 (Hedge first, premium second). Part II index.