Why your legal-AI tool pretends it read your chart

Verify it yourself — free, no login

See how AI medical-record review links every fact to the exact Bates page that proves it — click any citation and jump straight to the record.

See the 60-second demo →
April 22, 2026 · MedLegal AI Editorial · 7-min read

If you uploaded 100 medical-record PDFs to an AI tool last month, roughly 63 were probably never actually read.

Not "read poorly." Never read at all — zero characters of text extracted, with the AI generating output from the filename, the page count, and the model's prior on what a chart of that size usually contains.

On our own production telemetry, across a power-user account of 1,402 chart PDFs uploaded in one quarter, 883 — 63% — had zero extractable text on first pass. Image scans. Faxes of faxes. Old EMR exports "printed" to image. Every page looked like a medical record; a standard PDF text-layer read returned an empty string.

63% of PDF chart uploads in our telemetry have zero extractable text on first pass. Every one of them requires OCR to analyze.

Most legal-AI tools do not apply OCR by default. Some don't apply it ever. They accept the upload, read whatever text layer is there (often nothing), and generate plausible-looking output anyway. It's an industry-wide quiet failure mode.

Why so many medical-record PDFs have no text

Medicine runs on fax. Healthcare is one of the last US industries where fax is a core clinical communication channel — and fax output is an image, not text. Three common sources, in rough order of prevalence:

  1. Fax-to-PDF captures. A records clerk prints the chart, feeds it into a fax, the receiving end re-digitizes as an image PDF, and that file passes through several firms before reaching you. The final artifact is a stack of grayscale images.
  2. Scanner output from paper charts. Older providers, rural clinics, and nursing homes still keep paper. When litigation asks for records, a clerk runs them through a flatbed. Output is image-only unless someone deliberately runs OCR — and they usually don't.
  3. EMR "print-to-image" exports. Some older EMRs, especially in long-term care, route releases through image pipelines. Even modern Epic or Cerner exports sometimes lose the text layer when pushed through rasterizing release-of-information portals.

A native text PDF has an invisible text layer — copy-paste and search work, and the AI reads the layer directly. A scan PDF has no text layer. Copy-paste returns nothing. Unless you run OCR, the AI is working with a blank page.

What most competitor tools actually do

When a legal-AI tool receives a 400-page image-scan PDF and extracts zero text, there are three honest responses: refuse ("this file requires OCR, not supported"), run OCR and charge for the compute, or analyze partial text and flag the rest.

A fourth response is the most common on the market: generate output anyway. The tool sees a filename like Smith_John_ED_Records_2024-03-15.pdf and a 412-page count, then produces a summary that looks like an ED chart summary — chief complaint, triage vitals, assessment, plan. None of it came from reading the document. It came from the model's prior on what an ED chart usually looks like.

This is the same failure mode that produced the fabricated citations in Mata v. Avianca (S.D.N.Y. 2023). In Mata the fabrication was cases and quotes. In chart review, the fabrication is the chart content itself.

If your AI tool charges you for analysis but can't prove which pages it read, you don't have an AI tool. You have a text generator.

How to test your current stack in five minutes

Don't take anyone's word for this, including ours. The test:

  1. Find an image-only PDF. Any scanned discharge summary works. Open it in Preview or Acrobat and try to select a word — if the cursor selects the whole page as an image, it's image-only.
  2. Note a specific phrase on page 3 or later — a medication dose, a vital sign, an exact time. Something visually verifiable on the page, not in the filename.
  3. Upload to your legal-AI tool and ask for a summary.
  4. Ask the tool to quote that exact phrase — "Quote me the medication dose on page 3."

If the tool quotes a dose that isn't on the page — especially a plausible one it invented — you have your answer. The OCR isn't happening; the tool is generating from priors, not reading your chart. If it correctly refuses or correctly quotes the phrase, it's doing the right thing. Most do neither.

A concrete example: Smith v. Regional Medical Center

Our demo case — Smith v. Regional Medical Center — is a missed pulmonary embolism: a 42-year-old woman presents with pleuritic chest pain, is misdiagnosed as costochondritis, discharged, and goes into cardiac arrest at home four days later. The defendant produced six PDFs: ED triage, ED physician note, discharge instructions, radiology report, code-team run sheet, and autopsy.

Four of the six were image-only scans. Only the ED triage and the radiology report had native text layers. A tool that silently skipped OCR would describe only the triage vitals and the CT-angio read — the two documents whose causation is least useful to the plaintiff. The discharge instructions (premature-discharge argument), the code-team sheet (downtime-to-ROSC interval), and the autopsy (where the PE is definitively established) would all be invisible. A tool "analyzing" six PDFs without OCR quietly strips out the four documents the case actually depends on.

What Records Analyzer does differently

We built Records Analyzer because the failure mode above was happening on every other stack we tested.

On OCR error rates. Tesseract isn't perfect — handwritten nursing notes stay hard. We flag low-confidence passages visibly and keep the source image accessible so an attorney or LNC can eyeball the original. An imperfect OCR pass that's transparent about confidence is fundamentally different from no OCR at all dressed up as analysis.

The broader honesty point

Legal AI marketing has run ahead of the engineering. Vendors demo on clean, text-native PDFs because those demos look great. The PDFs your clients actually send are fax captures and scanner exports. If your current tool is doing the OCR work, great. What isn't fine is an industry-wide pattern of charging attorneys for analysis of documents that were never opened.

Test it on your own records

Records Analyzer is free to try — no credit card, three cases on the house. Run the five-minute test above on any image-only chart PDF. If the quoted phrase matches the page, we're doing the work. If it doesn't, we'd like to know.

Try Records Analyzer → · See the full platform

See the AI cite its source — no login
Most legal AI is wrong 17–33% of the time. Watch MedLegal AI pin every finding to the exact record page — click any citation and it jumps to the line that proves it.
Watch the 30-second demo →