You bought AI.

Why are you still paying people to check every answer?

Hammer Labs tests when healthcare agents are wrong, when they should refuse, and where human review still belongs.

See a real receipt → Start a six-week assessment →

See one answer, and the claim it refuses to make

Pick a billing code and an insurer. This answer was served from public payer evidence on SEP 6, 2026, copied unedited. Missing evidence returns no figure; a published rate never becomes a paid amount or a coverage decision by implication.

Basis
measured, from a rate extract of AUG 27, 2026 and fee schedules of AUG 21, 2026.
Not
The allowed amount, what was paid, or whether the payer covers the code.
Reproduce
Call rate_position with the same code against hammer-coverage-intelligence:2026-08-31+11.

Ask the live agent → Read all 25 refusals →

QUERY rate_position(code: "91200", payer: "aetna")
CPT 91200 Liver elastography, non-imaging VCTE (FibroScan)
MEDIAN$20.77
P25$13.66
P75$34.61
N ROWS11,859,263

N COUNTS · CURATED TREE One row per distinct rate, layered re-pulls removed. n is a count of contract rows in this extract, not a count of transactions or of providers who perform the priced service, and it has not passed a claims-joined ghost-rate or plausibility filter.

ANCHOR

Medicare Physician Fee Schedule, vintage 2026-08-21: 1.01 RVU.

DECLINED

Priced on the Physician Fee Schedule as relative value units, not dollars. Converting relative value units to a dollar figure needs the annual conversion factor and a locality-specific geographic adjustment, neither of which this world holds, so no ratio to the anchor is computed for this code.

RATE VINTAGE AUG 27, 2026 FEE SCHEDULE VINTAGE AUG 21, 2026 RECEIPT R11
SERVED ANSWER, UNEDITED. SOURCES: RATE EXTRACT AUG 27, 2026; CMS FEE SCHEDULE →

Build the answer key before asking the model

Most healthcare AI still ends with a nurse, coder or medical director checking every output. Real claims rarely carry a clean answer key. Hammer builds one from preserved evidence, then measures where an agent is wrong and where it should have refused. Payers now face decision deadlines and audit requirements under CMS-0057-F; review must become measurable, not disappear into another queue.

01 · PRESERVE

Keep the record that existed then

Capture each source with its date, basis and original artifact before upstream files are replaced.

02 · TEST

Plant errors with known answers

Run agents against records where expected decisions and failure conditions exist before any model responds.

03 · REFUSE

Stop unsupported sentences

Return unknown when evidence is missing. Name what would settle the question instead of laundering absence into confidence.

Public evidence proves the method. Private evidence runs the work.

PUBLIC · LIVE Free · rate limited

Healthcare coverage evidence

115 billing codes, 9 payers, Medicare fee schedules, coding edits and coverage articles. Install it, run a query and inspect its receipt. Public material demonstrates behavior; it is not the paid product. The CMS-derived archive behind it is free to read in the browser, dated vintages included.

Install the free plugin → Open the reader →

PRIVATE · ANNUAL from $150k/yr

Your claims, contracts and policies

Same tests, dated record and refusal boundaries on your evidence. Hosted by us or inside your VPC. Our team works alongside yours until your people run it. Our hours on your account go down every quarter. We show you the number.

Start with the assessment →

Start with one six-week assessment

Bring one code family or one cohort. We test where the agent is wrong, where it should refuse, and what measuring the same behavior on your claims requires.

Scope
One code family, up to three payers, one market, two dated versions.
Time
six weeks, fixed scope.
Price
$20k, credited against year one.
Your data
None required. No PHI.
Boundary
No payment prediction and no claim about what your contract says.

Apply for an assessment →

AFTER SIX WEEKS

Convert only if the platform boundary holds

Useful work becomes a private world on your evidence, priced annually. Bespoke findings become reusable tests, configurations and refusal classes. Work that cannot cross that boundary stays outside the platform rather than becoming permanent custom code.

Assessment fee comes off year one. Your team can run the system on our infrastructure or inside your VPC.

Proof before promise

25

Published refusal classes

Each names the unsupported sentence and, where known, evidence that would settle it.

15+

Years inside regulated data

Claims, payer and provider data, clinical infrastructure, lending, insurance and defence. Not domain expertise added after the product.

SEP 14, 2026

One buyer path, one proof and one platform boundary

Every release and preserved vintage gets a dated entry. Nothing planned appears as shipped.