What 13,000 Insurance Claim Documents Actually Look Like
96% were photographs of paper. The exact diagnosis was recoverable on one claim in ten. The whole book cost under $25 to read. What a real claims corpus teaches you about automating one.
Sep 26, 2026
Most writing about insurance claims automation starts from a clean sample. We started from a European insurer's entire claims book: just over 13,000 documents across some 6,500 claims, nine months of whatever policyholders actually sent in.
The first measurement reset every assumption after it.
What does a real claims book actually contain?
96% of it has no text layer at all. Not scanned PDFs with selectable text. Photographs. A clinic prints an invoice on paper, the owner photographs it on a phone, the photo is wrapped in a PDF and emailed. Median resolution 256 ppi, a third below 200, some of it blurred or cropped.
So the front door is not parsing. It is reading a picture of a page the way a person would, and everything downstream depends on doing that well.
Two smaller findings landed just as hard. Filenames said 54 documents contained both an invoice and a claim form; the content said 642. And searching the text for the word "facture" pulled in 4,859 claim forms, because the form's own printed boilerplate says "not accompanied by the corresponding invoice". Classify documents on content, never on filenames, and never on a single keyword.
How do you know the extraction is right if nobody labelled anything?
Find the checksum your domain already prints. An invoice states its own total. If the line amounts you pulled out add up to the total on the page, the reading is almost certainly correct.
That is a free, dense, automatic grade over every document in the book, available before anyone labels anything. It is the single most useful thing we did, and it decided every design question after it:
- A hand-tuned regex parser reconciled on 16% of invoices.
- The same model reading the OCR transcription reached 92%.
- The same model, same prompt, looking at the page image instead: 96%.
Reading the page beats reading a transcription of it, and the gap widens on longer invoices, which is the signature of a layout problem rather than a character-recognition one. OCR flattens a five-column table into a line of text and throws away which column a number sat in.
Most domains have a checksum like this and never use it. Invoices state totals. Ledgers balance. French business identifiers carry a Luhn check digit, which is why preferring the Luhn-valid candidate took our identifier accuracy from 62.6% to 91.6% with no model, no lookup and no cost.
Why read every document twice?
Because model confidence is not evidence. On the handwritten claim forms, we had two independent readers extract the stated diagnosis: one from the scan text, one from the page image. Where both named a condition, they contradicted each other on roughly a third of forms while both reported high confidence.
A single reader would have looked completely trustworthy and been wrong on a third of them. So the rule became: a reading counts when two readers who never saw each other's answer agree. Everything else is dropped, and where only one reader could make out the page, a third model from a different vendor arbitrates.
That is expensive in data. It cut a usable set roughly in half. It is the only reason the remainder can be trusted at all.
The corollary matters as much: abstention has to be a first-class output. On a claim document an invented diagnosis is worse than a blank one, because nothing downstream can tell a guess from a reading. Our condition classifier declines on 12% of claims rather than guess, and is wrong on 2%.
What did not work
Three things we expected to help, measured, and dropped.
Automated prompt optimisation gave nothing. We compiled the extraction prompt against the free checksum metric with DSPy, 60 training examples, proper held-out test set. Baseline 85.0%. Compiled 85.0%. A prompt carrying a week of failure analysis was already at a local optimum, and no amount of search improved it.
Standardising the wording made prediction worse. We built a vocabulary that folds more than 9,000 different clinic spellings into a single canonical name per product and procedure. Feeding those clean names to the classifier instead of the clinic's own raw wording dropped accuracy from 63.0% to 61.9%. Standardisation is for aggregation and comparison. The messy original carries detail a canonical label throws away.
A quarter of our errors were the taxonomy, not the model. Two categories in our own list overlapped, and fixing the wording of the definitions, changing nothing else, moved accuracy three points. Before blaming a model, check whether your categories are mutually exclusive.
What does it cost?
Under $25 for the entire book. Local OCR was free and ran on a laptop. Reading roughly 6,500 invoices into structured lines cost $12.68. Reading some 5,500 handwritten claim forms twice each cost $5.96. Classifying every claim came to $1.60.
The expensive resources were never compute. They were access to the corpus, the permissions, and the domain knowledge to know what a billed line is. Choose models on accuracy, not price, because at this scale price is not a real constraint.
The honest limit
Automating extraction is now straightforward. Automating interpretation is not.
We wanted to recover the exact condition behind a visit from the invoice alone. On this book it works on one claim in ten, and no model choice changes that: the median invoice is two lines of routine care, and a vaccination has no condition to recover. Asking a narrower question, placing the claim in one of seventeen categories rather than naming the pathology, gets 67%.
That gap is the real lesson. The pipeline was never the hard part. Knowing which question the documents can actually answer was.
If you are looking at a pile of documents and wondering what is recoverable from them, the answer is usually measurable in a day for the price of a coffee. We do this kind of document automation work as a fixed-scope slice, and we published what 30,000 messy records look like for the same reason: the corpus always tells you more than the plan did. If you want to know what is recoverable from yours, tell us what you have.