MedART is Compliance-Ready — Architecture designed to support HIPAA, GDPR & regional regulatory requirements.
Back to Blog

AI Scribe Accuracy: What Actually Goes Wrong in a Clinical Note

Clinicians fear the AI inventing things. The evidence says it mostly leaves things out — and omission is far harder to catch on review.

Prashant Talesara Prashant Talesara
September 3, 2026 7 min read

Ask a clinician why they have not adopted an AI scribe and the answer is rarely about speed. It is some version of: I don’t trust it to get the note right.

That instinct is correct. What is usually wrong is the picture behind it — because the failure clinicians brace for is not the failure the evidence actually finds.

What actually goes wrong is omission, not invention

The fear is hallucination: the AI writes something that was never said. It does happen. It is not, however, the main event.

A 2025 evaluation published in Mayo Clinic Proceedings: Digital Health assessed five ambient documentation platforms against fourteen simulated clinical encounters, classifying every error as omission, commission or partially correct (study). The findings are worth sitting with:

  • Omission accounted for 83% of identified errors. The note left something out far more often than it made something up.
  • Mean note error rate was 26.3% (95% CI 17.0–31.0).
  • 19.5% of transcript-level errors propagated into the finished note — most were caught somewhere in the pipeline, but a fifth were not.
  • An average of 3.0 errors per case carried potential for moderate-to-severe harm, with one case reaching 21.
  • Accuracy varied substantially between platforms and within the same platform.

Now consider what that means for a clinician reviewing a draft at the end of a clinic.

An invented detail announces itself. You read “patient reports no abdominal pain” for a patient who spent the consultation describing abdominal pain, and you catch it, because the sentence contradicts your memory of the room.

An omission does not announce itself. The note reads cleanly. Every sentence in it is true. Nothing is wrong on the page — the problem is a sentence that is not on the page, and catching it requires you to reconstruct the conversation from memory and notice an absence. Twenty consultations later, that reconstruction is not reliably happening.

This is why “we review every note” is a weaker control than it sounds when the dominant error is subtractive. Review catches commission well. It catches omission badly. Any serious answer to scribe accuracy has to be built around that asymmetry rather than around a reassurance that a human looks at it.

A separate narrative review of eighteen studies published between 2019 and mid-2025 reached a compatible conclusion: ambient scribes consistently reduce documentation burden, and they also produce frequent omissions alongside occasional clinically significant hallucinations (Cardiovascular Diagnosis and Therapy). Both things are true at once, and a clinic evaluating the technology has to hold both.

Why an IVF consult is a harder surface than general medicine

Published scribe evaluations are largely drawn from general ambulatory care. Fertility medicine is a harder case, for reasons that map directly onto what omission-type errors tend to hit.

The clinically decisive content is short and numeric. A fertility consultation turns on things said briefly and once: a dose adjustment, an oestradiol value, a cycle day, an antral follicle count. Long explanatory passages are relatively easy for a model to capture. Short unrepeated numbers are exactly the material most likely to be dropped — and in stimulation management, the number is the clinical decision.

Drug and protocol naming is inconsistent. The same molecule carries different brand names across markets, and protocols are referenced by shorthand that varies between clinics and even between clinicians in one clinic. A general-purpose speech model has no reliable prior for any of it.

Consultations frequently cross a language boundary. In the markets where IVF grows fastest, the clinician and patient often do not share a first language, and part of the consultation happens through a coordinator or family member. Every additional voice and accent widens the error surface.

Emotional register competes with clinical content. Fertility consultations carry difficult conversations, and the parts that matter clinically are interleaved with parts that matter enormously to the patient but are not documentation. Weighting between them is a judgement, and judgement is where models are weakest.

None of this argues against using a scribe in fertility care. It argues that a scribe trained on general ambulatory speech will underperform here, and that vocabulary coverage is a real differentiator rather than a spec-sheet line. MedAI Scribe is built around IVF vocabulary specifically — 320+ IVF-specific terms across 90+ languages — which narrows the surface but does not eliminate it. Nothing eliminates it.

Review-in-the-loop as architecture, not disclaimer

Most vendors will tell you a clinician should review every note. That statement costs nothing and proves nothing. The question worth asking is what the system does to make review real, and what it records when review happens.

There is a meaningful difference between three postures that all get described as “human in the loop”:

  • A disclaimer. The product tells you to check the output. Nothing in the software changes whether you did.
  • An approval click. The clinician confirms the note. The system stores a confirmation but not what was confirmed, so a note read carefully and a note approved in bulk are indistinguishable afterwards.
  • A governed review step. The draft has no clinical standing until signed, the system records what changed between draft and signature, and the signature is attributable and tamper-evident.

Only the third produces anything defensible. In MedAI Scribe, every AI note is a draft until the clinician signs it, edits are diff-tracked with timestamp and operator identity, and approved notes carry a tamper-evident audit hash, exportable in registry-ready formats for ESHRE, SART, DHA, MOH, KARM and PSRM.

The diff is the part clinics undervalue at purchase and rely on afterwards. It answers questions no approval click can:

  • What did the model produce, and what did the clinician change?
  • Which parts of a note were authored by the AI and which by a person?
  • Are edits concentrated in particular sections — suggesting a systematic weakness worth configuring around, rather than random error?
  • If a documentation question arises two years later, can you reconstruct who was responsible for which sentence?

That last question is the one that matters in a fertility clinic, where records are examined long after the cycle and where the chain of accountability across systems is already under scrutiny. A signed note with no edit history asserts that a clinician approved it. A signed note with a diff demonstrates what they approved.

There is also a quieter operational benefit. Once edits are tracked, the review step becomes measurable. A clinic can see whether a clinician who signs notes in four seconds has a different downstream correction rate than one who takes forty, and can address that as a training question rather than discovering it during an audit.

What the evidence means for your evaluation

Two adjustments to how most clinics assess this, both following from the findings above.

Ask about omission, not accuracy. A vendor accuracy percentage almost always describes transcription of clear speech. That is a real measurement and it is not the one that matters, because the errors that reach patients are things the finished note failed to carry. Ask instead: on a consultation we supply, what did the note miss? Then check the draft against the recording yourself rather than against your memory.

Test with your hardest consultation, not your cleanest. Platform variability in the published evaluation was substantial, including within a single platform across different encounters. A demo built on a clear, scripted, single-language consultation tells you close to nothing about a real one with a coordinator translating and a dose change mentioned in passing.

Beyond those two, the broader selection criteria — vocabulary training, EMR field population, deployment, data handling — are covered in what to ask before choosing an AI medical scribe, and the comparison against dictation and typing is in AI scribe vs dictation vs typing. This post is deliberately narrower: those cover whether and which, this covers whether you can trust the output.

The honest position

An AI scribe will produce notes with errors, most of them omissions, at a rate that varies by platform and by consultation. That is what the published evidence shows, and any vendor implying otherwise is describing a product that has not been measured the way these five were.

What makes the technology safe to use is not a model that never errs. It is a workflow where the draft has no standing until a clinician signs it, where what changed between draft and signature is recorded and attributable, and where the clinic reviews knowing that the likeliest error is something missing rather than something invented.

Trust is not a property of the model. It is a property of the process around it — and that, unlike the model, is something you can actually inspect before you buy.

See the review step, not the demo

Bring us a consult and watch the edit trail

A 30-minute session on a real consultation — what the draft captured, what it missed, and exactly what the audit record shows between draft and signature.

Topics

MedAI Scribe Clinical Documentation AI Governance Quality MedART
Prashant Talesara — Co-Founder, Meddilink EMR

Co-Founder, Meddilink EMR

Prashant Talesara is a co-founder of Meddilink EMR, the purpose-built IVF EMR platform. He is also Co-Founder & CTO at Datareel.ai — where he focuses on AI-powered hyper-personalization — and a Co-Founder at Kansoft. His work centers on building scalable technology that empowers industries, bringing engineering leadership and an AI-first approach to the products he helps create.

Frequently Asked Questions

How accurate are AI scribes in a clinical setting?
Accuracy varies widely between platforms and is not a single number. A 2025 evaluation of five ambient scribe platforms across simulated encounters, published in Mayo Clinic Proceedings: Digital Health, found a mean note error rate of 26.3% with substantial variability both between platforms and within the same platform. Vendor accuracy figures usually describe transcription of clear speech, which is a different and easier measurement than the correctness of the finished note.
Do AI scribes hallucinate in medical notes?
They can, but that is not the dominant failure mode. In the same evaluation, omission accounted for 83% of identified errors — the note left something out rather than inventing it. That matters for review, because an invented detail looks wrong on screen while a missing one looks like a complete note.
What does review-in-the-loop actually mean?
It means the AI output is a draft with no clinical standing until a clinician reviews and signs it, and that the system records what changed between draft and signature. In MedAI Scribe every note is a draft until the clinician signs it, edits are diff-tracked with timestamp and operator identity, and approved notes carry a tamper-evident audit hash. A disclaimer saying notes should be checked is not the same thing.
Why is IVF documentation harder for an AI scribe than general medicine?
Density and consequence. A fertility consult carries protocol names, drug names that differ by market, dose adjustments, cycle-day references and numeric results, often across a language boundary between clinician and patient. Much of that is short, numeric and unrepeated in the conversation, which is exactly the material most likely to be dropped.
Should a clinic wait until AI scribe accuracy improves?
Waiting for a fixed accuracy threshold misreads the problem. The published variability is between platforms and between note sections, not a single figure improving over time, and the control that makes a scribe safe today is the review step rather than the model. The practical question is not whether the draft is perfect but whether your workflow reliably catches what it missed.