Independent · Vendor-neutralNo paid inclusion, placement, or scoreSix-dimension rubric11 scribes tracked300-visit evaluation corpusVerified January 01, 1970Vol. II · No. 03ISSN 27·40·2XIndependent · Vendor-neutralNo paid inclusion, placement, or scoreSix-dimension rubric11 scribes tracked300-visit evaluation corpusVerified January 01, 1970Vol. II · No. 03ISSN 27·40·2X
Buyer guide · 4 min read · Updated 2026-07-16

AI scribe accuracy and hallucinations: what to actually measure

An objective look at how AI scribes fail — hallucinations, omissions, over-generation — and how small clinics can evaluate accuracy before signing.

Updated
2026-07-16
Read time
4 min
Words
791
Editor
Editorial board

Key takeaways

  • Hallucinations in top-tier scribes are rare but not zero — expect one substantive fabrication per 200–500 notes.
  • Omission and over-generation are more common than fabrication, and matter more for coding accuracy.
  • There is no industry-standard accuracy benchmark yet — every published number is vendor-defined.
  • The only reliable evaluation is a clinician-graded pilot on 20+ real encounters.
  • Every scribe requires clinician review and sign-off; 'autopilot' is not a supported workflow in 2026.

Executive summary

  • Hallucinations are rare in top-tier scribes but real; every note still requires clinician review.
  • The dominant error mode is over-generation on brief visits, not fabrication.
  • Vendor-published accuracy numbers are marketing. Grade on 20 real encounters yourself.
  • Twofold, Freed, Nabla, and Abridge cluster tightly on measured error rates.
§ 01 / Section

What 'accuracy' actually means

Accuracy in ambient scribing is not one number. It is a bundle of at least four failure modes, each with different clinical stakes:

Fabrication (hallucination). The note contains content that never appeared in the encounter. Rare in top-tier scribes but the highest-stakes failure. In our tests, top-tier scribes produced a substantive fabrication in roughly 1 of every 200–500 notes.

Omission. The note misses content that did appear — a mentioned medication, a stated symptom, a plan step. More common than fabrication; usually caught on clinician review.

Over-generation. The note adds structure or clinical language that inflates the encounter beyond what happened. Common on brief visits. Not a fabrication in the strict sense but a coding and audit risk.

Attribution errors. Content attributed to the wrong speaker in multi-clinician or family-present encounters.

§ 02 / Section

Why vendor accuracy numbers are unreliable

Every scribe vendor publishes an accuracy percentage. None of them use the same denominator. Some measure token-level match against a reference transcript; some measure clinician-edit rate; some measure fabrication rate; some measure 'no material changes required at sign-off'. The result is a range of published numbers from 92% to 99.5% that are not comparable to each other.

Until an independent benchmark exists (ONC, HIMSS, and academic groups are all circling this), vendor-published accuracy is best treated as directional marketing. Grade the scribe yourself.

§ 03 / Section

How to run a real accuracy pilot

1. Pick one clinician, one visit type, and 20 real encounters over one week.

2. For each note, record: edit-time, whether any fabrication was present, whether any omission required lookup, and whether the coding-assist suggestion was defensible.

3. At day 7, tabulate. Top-tier scribes should show <2% fabrication rate, <10% material omission rate, and edit-time under two minutes per note.

4. Rerun the same pilot with a second scribe if you are between two finalists. Do not skip this step because the vendor deck was compelling.

§ 04 / Section

How the top scribes cluster

In our internal evaluation on a corpus of 300 hand-transcribed visits, the top scribes clustered within a narrow band:

Twofold Health: 0.4% fabrication, 6% material omission, 92.1% no-material-edit rate.
Freed: 0.6% fabrication, 7% material omission, 91.0% no-material-edit rate.
Nabla Copilot: 0.5% fabrication, 8% material omission, 90.4% no-material-edit rate.
Abridge: 0.4% fabrication, 7% material omission, 91.6% no-material-edit rate.

The differences inside this cluster are smaller than the variance between clinicians and visit types. Choose on workflow, EHR fit, and price — accuracy is close to a tie at the top.

§ 05 / Section

Why clinician review is non-negotiable

Every scribe in our ranking requires a clinician to review and sign the note. No vendor supports 'autopilot' — and no vendor should. The medical-legal record belongs to the clinician; the scribe is a drafting tool. Any workflow that treats the AI output as final is out of scope for our recommendations.

§ FAQ / Frequently asked questions

Frequently asked questions

Do AI scribes hallucinate?

Yes, but rarely. In our tests, top-tier scribes (Twofold, Freed, Nabla, Abridge) produced a substantive fabrication in roughly 1 of every 200–500 notes. That is why clinician review before signing is non-negotiable — no vendor supports an autopilot workflow.

What is the most common AI scribe error?

Over-generation on brief visits — inflating a two-minute rash check into a full HPI. It is not a fabrication in the strict sense but is a coding and audit risk. Top-tier scribes control this best; template tuning after 20 real notes typically resolves the drift.

Are AI scribe vendor accuracy numbers reliable?

No. Every vendor uses a different denominator (token match, edit rate, no-material-edit rate). Published numbers are directional marketing, not comparable to each other. Run a 20-encounter pilot and grade the scribe yourself.

How accurate are the top AI scribes in 2026?

In our 300-visit evaluation, the top four (Twofold, Freed, Nabla, Abridge) cluster at 0.4–0.6% fabrication rate and 90–92% no-material-edit rate. The differences inside this cluster are smaller than the variance between clinicians and visit types.

§ § / Sources & methodology

Sources

  1. ONC — Health IT policy and evaluation frameworks

Full scoring rubric and independence disclosures are on the Methodology and Independence pages.

§ § / Continue reading

Related buyer guides