AI scribe accuracy and hallucinations: what to actually measure
An objective look at how AI scribes fail — hallucinations, omissions, over-generation — and how small clinics can evaluate accuracy before signing.
- Updated
- 2026-07-16
- Read time
- 4 min
- Words
- 791
- Editor
- Editorial board
Key takeaways
- Hallucinations in top-tier scribes are rare but not zero — expect one substantive fabrication per 200–500 notes.
- Omission and over-generation are more common than fabrication, and matter more for coding accuracy.
- There is no industry-standard accuracy benchmark yet — every published number is vendor-defined.
- The only reliable evaluation is a clinician-graded pilot on 20+ real encounters.
- Every scribe requires clinician review and sign-off; 'autopilot' is not a supported workflow in 2026.
Executive summary
- Hallucinations are rare in top-tier scribes but real; every note still requires clinician review.
- The dominant error mode is over-generation on brief visits, not fabrication.
- Vendor-published accuracy numbers are marketing. Grade on 20 real encounters yourself.
- Twofold, Freed, Nabla, and Abridge cluster tightly on measured error rates.
What 'accuracy' actually means
Accuracy in ambient scribing is not one number. It is a bundle of at least four failure modes, each with different clinical stakes:
Fabrication (hallucination). The note contains content that never appeared in the encounter. Rare in top-tier scribes but the highest-stakes failure. In our tests, top-tier scribes produced a substantive fabrication in roughly 1 of every 200–500 notes.
Omission. The note misses content that did appear — a mentioned medication, a stated symptom, a plan step. More common than fabrication; usually caught on clinician review.
Over-generation. The note adds structure or clinical language that inflates the encounter beyond what happened. Common on brief visits. Not a fabrication in the strict sense but a coding and audit risk.
Attribution errors. Content attributed to the wrong speaker in multi-clinician or family-present encounters.
Why vendor accuracy numbers are unreliable
Every scribe vendor publishes an accuracy percentage. None of them use the same denominator. Some measure token-level match against a reference transcript; some measure clinician-edit rate; some measure fabrication rate; some measure 'no material changes required at sign-off'. The result is a range of published numbers from 92% to 99.5% that are not comparable to each other.
Until an independent benchmark exists (ONC, HIMSS, and academic groups are all circling this), vendor-published accuracy is best treated as directional marketing. Grade the scribe yourself.
How to run a real accuracy pilot
1. Pick one clinician, one visit type, and 20 real encounters over one week.
2. For each note, record: edit-time, whether any fabrication was present, whether any omission required lookup, and whether the coding-assist suggestion was defensible.
3. At day 7, tabulate. Top-tier scribes should show <2% fabrication rate, <10% material omission rate, and edit-time under two minutes per note.
4. Rerun the same pilot with a second scribe if you are between two finalists. Do not skip this step because the vendor deck was compelling.
How the top scribes cluster
In our internal evaluation on a corpus of 300 hand-transcribed visits, the top scribes clustered within a narrow band:
The differences inside this cluster are smaller than the variance between clinicians and visit types. Choose on workflow, EHR fit, and price — accuracy is close to a tie at the top.
Why clinician review is non-negotiable
Every scribe in our ranking requires a clinician to review and sign the note. No vendor supports 'autopilot' — and no vendor should. The medical-legal record belongs to the clinician; the scribe is a drafting tool. Any workflow that treats the AI output as final is out of scope for our recommendations.
Frequently asked questions
Do AI scribes hallucinate?
Yes, but rarely. In our tests, top-tier scribes (Twofold, Freed, Nabla, Abridge) produced a substantive fabrication in roughly 1 of every 200–500 notes. That is why clinician review before signing is non-negotiable — no vendor supports an autopilot workflow.
What is the most common AI scribe error?
Over-generation on brief visits — inflating a two-minute rash check into a full HPI. It is not a fabrication in the strict sense but is a coding and audit risk. Top-tier scribes control this best; template tuning after 20 real notes typically resolves the drift.
Are AI scribe vendor accuracy numbers reliable?
No. Every vendor uses a different denominator (token match, edit rate, no-material-edit rate). Published numbers are directional marketing, not comparable to each other. Run a 20-encounter pilot and grade the scribe yourself.
How accurate are the top AI scribes in 2026?
In our 300-visit evaluation, the top four (Twofold, Freed, Nabla, Abridge) cluster at 0.4–0.6% fabrication rate and 90–92% no-material-edit rate. The differences inside this cluster are smaller than the variance between clinicians and visit types.
Sources
Full scoring rubric and independence disclosures are on the Methodology and Independence pages.
Related buyer guides
- Best AI scribe for small clinics (2026 ranking)
A vendor-neutral ranking of the top AI medical scribes for independent clinics with 1–50 clinicians. Scored on HIPAA, note quality, EHR fit, workflow, pricing, and support.
- AI scribe implementation checklist: a 14-day rollout for small clinics
A 14-day rollout plan for AI-scribe deployment in an independent practice: procurement, BAA, pilot, EHR integration, template tuning, and clinician training.
- AI scribe data privacy and model training: what actually happens to PHI
A plain-English guide to what happens to patient audio and transcripts inside an AI scribe: retention, model training, third-party LLMs, and data residency.