AI scribe data privacy and model training: what actually happens to PHI
A plain-English guide to what happens to patient audio and transcripts inside an AI scribe: retention, model training, third-party LLMs, and data residency.
- Updated
- 2026-07-16
- Read time
- 4 min
- Words
- 792
- Editor
- Editorial board
Key takeaways
- Every top-tier scribe contractually forbids training on identified PHI; ask for the clause number.
- De-identified aggregate training is common and legal — decide whether your practice is comfortable with it.
- Third-party LLM providers (OpenAI, Anthropic, Google) must sit under BAA in the data chain.
- Default audio retention varies from minutes to 90 days; retention is almost always configurable.
- US data residency is standard among ranked vendors; ask before assuming.
Executive summary
- Every scribe in our ranking forbids training on identified PHI in its BAA.
- De-identified aggregate training is common and legal — practices should decide whether that is acceptable.
- Third-party LLM providers must be under BAA; ask which providers and how.
- Retention defaults vary; configure them explicitly during procurement.
The four data flows to check
Audio in transit. Encrypted (TLS 1.2 or 1.3) across all ranked vendors. Table stakes.
Audio at rest. Encrypted (AES-256 at rest across all ranked vendors). Retention varies — anywhere from immediate deletion after note generation to 90 days by default.
Transcript / note storage. Encrypted, but retention is typically longer than audio because the note is the deliverable. Confirm whether your EHR is the system of record after write-back.
Third-party model calls. The scribe likely uses at least one external LLM (OpenAI, Anthropic, Google). That provider must be a downstream business associate under BAA, and its enterprise terms must forbid training on PHI.
Identified vs de-identified training
Every top-tier scribe forbids training on identified PHI. This is universal and non-negotiable — ask to see the exact BAA clause.
De-identified aggregate training (using de-identified transcripts to improve models) is a different question. It is legal under HIPAA's Safe Harbor de-identification standard (45 CFR 164.514(b)(2)), and most vendors do it. Some let practices opt out at contract time. Whether it is acceptable is a practice-level policy call, not a compliance requirement.
Ask the vendor: 'Do you use de-identified transcripts to improve your models? Can we opt out? What is the exact de-identification method — Safe Harbor or Expert Determination?'
Which LLMs are in the chain
Most ambient scribes route to at least one third-party LLM. Common configurations in 2026:
Ask which providers are in the chain, whether each is under BAA, and whether the vendor will notify you before adding a new one.
Retention defaults to check
Data residency
All scribes in our 2026 ranking store PHI in US regions by default. If the vendor uses a global cloud provider, ask which region and whether data is ever routed outside the US for processing (some multi-model routing does this — Twofold, Freed, and Nabla confirm US-only routing).
Frequently asked questions
Do AI scribes train on my patient data?
Every scribe in our 2026 ranking contractually forbids training on identified PHI in its BAA. De-identified aggregate training (using Safe-Harbor-de-identified transcripts to improve models) is common and legal, and most vendors do it by default. Ask whether you can opt out at contract time.
Which third-party LLMs do AI scribes use?
Most route to at least one external model, typically Azure OpenAI under BAA. Twofold, Freed, Heidi, and Nabla all disclose their model chain on request. Abridge runs a substantially in-house stack. Ask which providers are in the chain and whether each is under BAA.
How long is patient audio kept?
Defaults vary from minutes (audio deleted after note generation) to 90 days. Almost every vendor allows the practice to configure retention. Confirm the default in writing during procurement.
Is my scribe data stored in the US?
All scribes in our 2026 ranking store PHI in US regions by default. Some multi-model routing sends processing outside the US unless explicitly configured; confirm US-only routing during procurement.
What is Safe Harbor de-identification?
The HIPAA Safe Harbor method (45 CFR 164.514(b)(2)) strips 18 specified identifiers from a record. A properly Safe-Harbor-de-identified transcript is no longer PHI and can be used to train models without BAA constraints. Ask your vendor whether Safe Harbor or Expert Determination is used.
Sources
Full scoring rubric and independence disclosures are on the Methodology and Independence pages.
Related buyer guides
- HIPAA guide for AI medical scribes (plain-English, 2026)
A plain-English guide to HIPAA compliance for AI scribes used by small clinics: BAA, encryption, retention, audit logging, breach obligations.
- AI scribe accuracy and hallucinations: what to actually measure
An objective look at how AI scribes fail — hallucinations, omissions, over-generation — and how small clinics can evaluate accuracy before signing.
- AI scribes for behavioral health and psychiatry clinics (2026)
How AI medical scribes handle psychotherapy, med-management, and intake visits in small behavioral-health practices — note formats, HIPAA specifics, and vendor fit.