Hammer Labs HL-P-006 SHEET 6 OF 6
HL-P-006 JUL 30, 2026 12 MIN READ 2,442 WORDS Hello World Models

Same Rule, Opposite Signs

Banking ran our definition-change test for us, better documented than we could have managed, and the answer was not the one anybody expects. On one day in January 2020 a single accounting rule moved reserves up at three banks and down at a fourth, and up and down inside the same bank. That gives a coherence test costing one group-by, and it also breaks a category we had been treating as one thing.

The previous post argued that regulated industries are the best place to build world models, because the regulator publishes the interventions you need to test one. It proposed a test to run first: find a date when a measurement definition changed, roll your model across it, and check that it moves the map without moving the territory.

Then we went looking for the same structure in a completely different industry, to find out which parts of the argument were about regulation and which parts were about healthcare. We picked US bank capital and credit-loss regulation, roughly 2016 to today.

It has all of it. Zero-parameter transitions written into statute. Published shocks with a legal deadline for publishing them. A definition change on a known date. Colliding clocks. A feedback loop so strong it makes the healthcare version look mild.

It also came back with a correction, and the correction is the good part.

The test we proposed, run properly, by somebody else

On 1 January 2020 a new accounting standard took effect for large US banks. CECL changed how a bank calculates the money it sets aside for loans that might go bad. The old rule: set money aside when a loss has probably already happened. The new rule: set money aside now for every loss you expect over the entire remaining life of the loan.

Same loans. Same borrowers. Same morning. Different number.

The regulators were unusually explicit that only the map had moved. The US Treasury, in a study Congress required:

“CECL does not change the cumulative amount of losses ultimately charged off from an asset, relative to ILM; it changes the timing of when those losses are recognized.”

The banking agencies went further and did something no health regulator did for us: they published a correction schedule. The stated purpose of the capital phase-in was to make the near-term effect of CECL “roughly comparable to the regulatory capital impact under the incurred loss methodology.” That is a regulator publishing the amount by which the map moved when the territory did not, with a dated schedule for unwinding it.

So: a definition change, a known date, an official statement that nothing real happened, and a published correction. It is the cleanest instance of our test that exists, and we did not have to run it.

Here is what it did.

Day-one change in loan loss allowance on 1 January 2020, when CECL took effect. JPMorgan, Citigroup and Bank of America all rose. Wells Fargo fell by $1.3bn.

Three banks reserved more. One reserved less. Same rule, same day, opposite signs.

That alone would be a good story. The part that changed how we think about the test is what happens when you look inside a single bank.

Day-one CECL change by portfolio. JPMorgan's card book rose $5.5bn while its wholesale book fell $1.6bn. Citigroup's consumer book rose $4.9bn while its corporate book fell $0.8bn. Wells Fargo's commercial book fell $2.9bn while card and auto rose.

One rule moved different parts of the same bank’s book in opposite directions on the same morning. And the sign is not mysterious. Lifetime expected loss reserves more against a credit card balance that will sit there for years, and less against a short commercial loan that matures before the losses would have arrived.

The sign of the move is a fact about how the subject is composed relative to the formula. It is not a fact about the world.

Which gives you a test that costs one group-by

That observation turns a story into something you can run:

When a measurement definition changes on a known date, group the change by subject and look at the signs. If every subject moved the same way, something real may have happened. If subjects moved in both directions, it is definitionally a remapping, and a model predicting a real-world change at that boundary has confused the map with the territory.

It needs no ground truth. It cannot be passed by training longer. It works at any sample size. It costs one group-by.

Try it on the actual numbers:

The sign test on the published CECL figures. The third series is simulated, and shows what the same test looks like when something real did happen.

We nearly got this exactly backwards. The sentence we would have written without checking is “CECL raised loan loss reserves.” It is false as stated, and it is false for a quarter of the banks that mattered. The direction always feels obvious in advance.

Note which error that is. Our recurring failure has been reading a difference in a label as a difference in the world. This is the mirror image: assuming a difference in the world from a label change, with the sign taken for granted. Same root, opposite direction.

The correction: not every published shock is a test

Here is where banking made us fix something.

Every February, the Federal Reserve publishes an imaginary recession in numeric detail and every large bank has to say what it would do to them. The 2026 scenario has unemployment reaching 10 percent, equities down 58 percent, commercial property down 39 percent. The deadline for publishing it is written into the regulation, 12 CFR 252.44, “no later than February 15.”

Then the Fed publishes its own answer, and within fifteen days every bank must publish its answer to the same question. A regulator wrote a cross-source disagreement test into a rule and put a deadline on both sides of it. We have nothing that good.

But unemployment does not then go to 10 percent. The scenario is a published counterfactual, not a published intervention.

The previous post listed phase-ins, partial mandates, calendar transitions and definition changes as “free detours” and proposed assembling them into a public held-out suite. That category needs splitting in two:

Both are useful. They are not the same kind of object, and a benchmark that scores them together reports one number built on two different epistemic things. Better to fix that before building the suite than after.

The bit where we have the better hand

Two more things banking cannot do, which are worth stating plainly because they are arguments for the harder domain rather than against it.

Every banking transition fires on the same day for everyone. CECL: 1 January 2020, for every calendar-year filer at once. The benchmark replacement: one date, every covered contract at once. The stress scenario: one scenario, all 32 banks, the same February. So every before-and-after comparison is confounded with whatever else happened on that date, and in CECL’s case, what else happened was a pandemic. The national emergency was declared on 13 March 2020, inside the same quarter. Two causes, one effect, no control group, and you get exactly one draw.

Healthcare’s clocks run per person. A coordination period starts on the individual’s own eligibility date. A post-transplant clock starts on the individual’s own transplant date. The same rule fires on thousands of different days across the population, staggered by construction, with calendar noise averaging out. That is a structurally better identification setting, and it is available to anyone willing to model at the level of a person rather than an institution.

Neither banking number is an observation. The bank computes its own projection; the Fed computes its own version of the same projection. Genuinely two models of one quantity, published a fortnight apart, and for several years the two were running on different definitions of the same line item, because the Fed’s supervisory models had not adopted CECL while the banks had. Anyone comparing the two numbers without knowing that was measuring the definitional gap and calling it disagreement.

Both are models of a hypothetical. In health, a payer’s filed rates and a hospital’s filed charges are two independent reports of the same real contracted price, filed by the two parties to it, under two different federal rules. That is the same disagreement test pointed at a quantity that actually exists.

And the feedback loop, which is worse there

The last post said the world reads your model. Banking has the strongest instance of that anywhere.

Before it was replaced, LIBOR was produced by asking a panel of banks a question, and the wording matters:

“At what rate could you borrow funds, were you to do so…?”

Could, and were you to do so. A question about something that did not happen, answered by the parties whose payments depended on the answer.

Three-month USD LIBOR: under $1 billion a day of observable trading, drawn to scale against the $223 trillion of contracts priced off it.

The manipulation scandals are the vivid part and the weaker argument. The structural indictment stands without any fraud at all: a benchmark that asks what rate you could borrow at has nothing to check itself against when almost nobody is borrowing. The fix was to stop asking and start counting. SOFR is computed from about $3 trillion a day of actual transactions.

And the model-monoculture warning is not ours either. Ben Bernanke, as Fed chairman, in 2013:

firms “would see a declining benefit to maintaining independent risk-management systems and would just adopt supervisory models instead”

He named it a model monoculture susceptible to a single common failure. Under the current rule the supervisory model’s output is a bank’s capital requirement, so the incentive to converge on it is not hypothetical. It has since been formalised: disclosure restores market confidence but misclassifies some healthy banks as risky, pushing banks toward portfolios that look safe under the regulator’s model. Optimal disclosure turns out to be non-monotone, more transparency is not always better.

What we are changing because of this

  1. A definition-change register, with a sign column. Every change gets its date, its affected population, the regulator’s own statement about whether the underlying moved, any published correction schedule, and the column banking taught us: did the number move in both directions?
  2. Split the detour suite in two. Realised interventions get scored. Published hypotheticals get used as a shared harness and never contribute to a score.
  3. Treat the staggered clock as the asset it is. Per-person timing is not a limitation of the harder domain. It is the reason the harder domain can score a detour at all.

One thing worth stealing outright, from an industry we did not even choose. A 2017 EU regulation required the Commission to publish an executable simulator so results measured under a new vehicle test could be converted back to the old scale. It ships a release every September, and it includes an anti-gaming mechanism: each run returns a random integer from 1 to 100, and if it lands in the top ten the vehicle gets pulled for real physical measurement.

That is a regulator shipping the map-to-map conversion as versioned software with a built-in audit sample. It is the answer to what a definition-change register should eventually become: not a document. A program, with a version number and a random audit.

The only two questions that matter

What does it cost me when this goes wrong? You book a real-world deterioration that did not happen, on the day a formula changed, and you act on it. Somebody’s reserves, staffing or approvals move for a reason that exists only in the measurement. The test that would have caught it is one group-by, run once, on a date the regulator published in advance.

Can I defend it to a regulator? This is the part that gets better rather than worse. The sign test produces an artifact an examiner can read without trusting your model at all: here is the date, here is the population, here is the direction each subgroup moved, here is why the mixed sign means the world did not change. No mechanism required. Specification, again.


Part of an ongoing series on building AI for regulated decisions you can explain, reproduce, and defend. Start at Hello World Models, then The Rules Are Half the Physics.


Sources and further reading

The definition change

The published shock

The benchmark

Model monoculture

The simulator worth stealing