Ask a clinician why they have not adopted an AI scribe and the answer is rarely about speed. It is some version of: I don’t trust it to get the note right.
That instinct is correct. What is usually wrong is the picture behind it — because the failure clinicians brace for is not the failure the evidence actually finds.
What actually goes wrong is omission, not invention
The fear is hallucination: the AI writes something that was never said. It does happen. It is not, however, the main event.
A 2025 evaluation published in Mayo Clinic Proceedings: Digital Health assessed five ambient documentation platforms against fourteen simulated clinical encounters, classifying every error as omission, commission or partially correct (study). The findings are worth sitting with:
- Omission accounted for 83% of identified errors. The note left something out far more often than it made something up.
- Mean note error rate was 26.3% (95% CI 17.0–31.0).
- 19.5% of transcript-level errors propagated into the finished note — most were caught somewhere in the pipeline, but a fifth were not.
- An average of 3.0 errors per case carried potential for moderate-to-severe harm, with one case reaching 21.
- Accuracy varied substantially between platforms and within the same platform.
Now consider what that means for a clinician reviewing a draft at the end of a clinic.
An invented detail announces itself. You read “patient reports no abdominal pain” for a patient who spent the consultation describing abdominal pain, and you catch it, because the sentence contradicts your memory of the room.
An omission does not announce itself. The note reads cleanly. Every sentence in it is true. Nothing is wrong on the page — the problem is a sentence that is not on the page, and catching it requires you to reconstruct the conversation from memory and notice an absence. Twenty consultations later, that reconstruction is not reliably happening.
This is why “we review every note” is a weaker control than it sounds when the dominant error is subtractive. Review catches commission well. It catches omission badly. Any serious answer to scribe accuracy has to be built around that asymmetry rather than around a reassurance that a human looks at it.
A separate narrative review of eighteen studies published between 2019 and mid-2025 reached a compatible conclusion: ambient scribes consistently reduce documentation burden, and they also produce frequent omissions alongside occasional clinically significant hallucinations (Cardiovascular Diagnosis and Therapy). Both things are true at once, and a clinic evaluating the technology has to hold both.
Why an IVF consult is a harder surface than general medicine
Published scribe evaluations are largely drawn from general ambulatory care. Fertility medicine is a harder case, for reasons that map directly onto what omission-type errors tend to hit.
The clinically decisive content is short and numeric. A fertility consultation turns on things said briefly and once: a dose adjustment, an oestradiol value, a cycle day, an antral follicle count. Long explanatory passages are relatively easy for a model to capture. Short unrepeated numbers are exactly the material most likely to be dropped — and in stimulation management, the number is the clinical decision.
Drug and protocol naming is inconsistent. The same molecule carries different brand names across markets, and protocols are referenced by shorthand that varies between clinics and even between clinicians in one clinic. A general-purpose speech model has no reliable prior for any of it.
Consultations frequently cross a language boundary. In the markets where IVF grows fastest, the clinician and patient often do not share a first language, and part of the consultation happens through a coordinator or family member. Every additional voice and accent widens the error surface.
Emotional register competes with clinical content. Fertility consultations carry difficult conversations, and the parts that matter clinically are interleaved with parts that matter enormously to the patient but are not documentation. Weighting between them is a judgement, and judgement is where models are weakest.
None of this argues against using a scribe in fertility care. It argues that a scribe trained on general ambulatory speech will underperform here, and that vocabulary coverage is a real differentiator rather than a spec-sheet line. MedAI Scribe is built around IVF vocabulary specifically — 320+ IVF-specific terms across 90+ languages — which narrows the surface but does not eliminate it. Nothing eliminates it.
Review-in-the-loop as architecture, not disclaimer
Most vendors will tell you a clinician should review every note. That statement costs nothing and proves nothing. The question worth asking is what the system does to make review real, and what it records when review happens.
There is a meaningful difference between three postures that all get described as “human in the loop”:
- A disclaimer. The product tells you to check the output. Nothing in the software changes whether you did.
- An approval click. The clinician confirms the note. The system stores a confirmation but not what was confirmed, so a note read carefully and a note approved in bulk are indistinguishable afterwards.
- A governed review step. The draft has no clinical standing until signed, the system records what changed between draft and signature, and the signature is attributable and tamper-evident.
Only the third produces anything defensible. In MedAI Scribe, every AI note is a draft until the clinician signs it, edits are diff-tracked with timestamp and operator identity, and approved notes carry a tamper-evident audit hash, exportable in registry-ready formats for ESHRE, SART, DHA, MOH, KARM and PSRM.
The diff is the part clinics undervalue at purchase and rely on afterwards. It answers questions no approval click can:
- What did the model produce, and what did the clinician change?
- Which parts of a note were authored by the AI and which by a person?
- Are edits concentrated in particular sections — suggesting a systematic weakness worth configuring around, rather than random error?
- If a documentation question arises two years later, can you reconstruct who was responsible for which sentence?
That last question is the one that matters in a fertility clinic, where records are examined long after the cycle and where the chain of accountability across systems is already under scrutiny. A signed note with no edit history asserts that a clinician approved it. A signed note with a diff demonstrates what they approved.
There is also a quieter operational benefit. Once edits are tracked, the review step becomes measurable. A clinic can see whether a clinician who signs notes in four seconds has a different downstream correction rate than one who takes forty, and can address that as a training question rather than discovering it during an audit.
What the evidence means for your evaluation
Two adjustments to how most clinics assess this, both following from the findings above.
Ask about omission, not accuracy. A vendor accuracy percentage almost always describes transcription of clear speech. That is a real measurement and it is not the one that matters, because the errors that reach patients are things the finished note failed to carry. Ask instead: on a consultation we supply, what did the note miss? Then check the draft against the recording yourself rather than against your memory.
Test with your hardest consultation, not your cleanest. Platform variability in the published evaluation was substantial, including within a single platform across different encounters. A demo built on a clear, scripted, single-language consultation tells you close to nothing about a real one with a coordinator translating and a dose change mentioned in passing.
Beyond those two, the broader selection criteria — vocabulary training, EMR field population, deployment, data handling — are covered in what to ask before choosing an AI medical scribe, and the comparison against dictation and typing is in AI scribe vs dictation vs typing. This post is deliberately narrower: those cover whether and which, this covers whether you can trust the output.
The honest position
An AI scribe will produce notes with errors, most of them omissions, at a rate that varies by platform and by consultation. That is what the published evidence shows, and any vendor implying otherwise is describing a product that has not been measured the way these five were.
What makes the technology safe to use is not a model that never errs. It is a workflow where the draft has no standing until a clinician signs it, where what changed between draft and signature is recorded and attributable, and where the clinic reviews knowing that the likeliest error is something missing rather than something invented.
Trust is not a property of the model. It is a property of the process around it — and that, unlike the model, is something you can actually inspect before you buy.
Topics
Co-Founder, Meddilink EMR
Prashant Talesara is a co-founder of Meddilink EMR, the purpose-built IVF EMR platform. He is also Co-Founder & CTO at Datareel.ai — where he focuses on AI-powered hyper-personalization — and a Co-Founder at Kansoft. His work centers on building scalable technology that empowers industries, bringing engineering leadership and an AI-first approach to the products he helps create.