AI in Healthcare

AI scribe hallucinations: why every fact needs a source

Accuracy percentages tell you little about whether a note is safe to sign. Traceability does: every fact linked to the words and audio it came from.

AI scribe accuracy is not one number. A note can be almost perfect and still contain the one invented finding that matters. The practical answer is traceability: every fact in the note should link to the exact words, speaker and audio it came from, and anything without a source should stay out of the note. That is the principle behind Verto's scribe.

What a “hallucination” looks like in a clinical note

Language models write fluent text. When the transcript is thin, they can fill the gap with what usually appears in a note for that presentation. In a clinic, that produces errors such as:

  • Invented exam findings. “Chest clear to auscultation” when no chest exam was done.
  • Template defaults. A normal review of systems filled in because the template expects one.
  • Negation flips. “Reports dizziness” when the patient said they had none.
  • Merged histories. A parent's or companion's condition recorded as the patient's.
  • Number drift. A dose, frequency, side or tooth number that changed between speech and note.
  • Confident plans. A follow-up interval or referral that was discussed as an option, recorded as decided.

Omissions are the quieter twin of hallucinations. A symptom mentioned once, in passing, can drop out of a summary. Both matter, because the signed note becomes the medical record, the basis for the next clinician's decisions and the evidence behind the claim.

Why accuracy percentages can mislead

You will see accuracy figures in scribe marketing. Treat them with care, including any you hear from us. A single percentage hides several questions:

QuestionWhy it changes the number
Accuracy of what?Word-level transcription accuracy is different from the accuracy of clinical facts in the final note
On which recordings?Clean, scripted English audio scores higher than real, bilingual, noisy clinic rooms
How were errors weighted?A misspelled word and a wrong dose are not the same error
Who judged it?Vendor staff, independent clinicians or an automated metric
Before or after review?What reaches the signed note depends on the review design as much as the model

That is why we do not publish an accuracy percentage for Verto. What matters is whether you can catch the errors that do happen, quickly, before you sign. Test any scribe on your own recordings, as described in our AI scribe evaluation checklist.

Traceability: the chain from audio to note

A traceable scribe keeps a link at every step of the pipeline:

  1. Audio segment. The moment in the recording where something was said.
  2. Utterance. The transcribed words, with the speaker and time stamp.
  3. Fact. The clinical statement extracted from those words, such as a symptom, finding, medication or plan.
  4. Note line. The sentence in the note built from that fact.
Fig. 01 · Process

The proof chain

  1. Audio segmentThe moment in the recording
  2. UtteranceThe words, with speaker and time
  3. FactThe clinical statement extracted
  4. Note lineThe sentence built from that fact
A traceable scribe keeps a link at every step, from the recording to the signed note.

If the chain is complete, you can go from any sentence in the note back to the audio in one step. If a sentence has no chain, the model wrote it without evidence, and it should not be in the record until a clinician confirms it.

How Verto enforces it

  • Tap-to-prove. Every fact card opens an evidence view with the exact utterance highlighted, one line of context, who said it and when. Once the visit is saved, you can replay just that audio segment, if the patient consented to audio retention.
  • Unsourced facts are held back. Facts Verto cannot trace to the visit go to a separate “needs a source” lane. They never reach the note unless a clinician confirms a source or edits them.
  • Certainty is visible. Fact cards show whether a fact is confirmed, likely or a negated finding (“Denies…”), so negatives are easy to check.
  • Accept only with a source. The accept action is only available when the fact has a source.
  • Rejections flow everywhere. Rejecting a fact rebuilds the note. The chart's assessment and plan follow the note, and exports are always rebuilt from reviewed facts, so a rejected fact cannot reach a printed or downloaded note.
  • Your edits are protected. If you typed into a section yourself, Verto does not overwrite it silently. It tells you the note is out of date and lets you choose when to rebuild.
  • Numbers in Arabic instructions are checked. When the plan is turned into Arabic for the patient, an automatic check confirms every number and dose unit survived the translation.

Staged orders follow the same rule. Prescriptions, labs, imaging and referrals are shown with the transcript excerpt that triggered them, and nothing is written to the EMR until you save.

How traceability changes the review

Traceability is not only a safety feature. It changes how long review takes, because clinicians stop re-reading whole notes and start checking single facts.

Proofreading a finished noteReviewing traceable facts
Unit of reviewA paragraphOne fact at a time
Checking a doubtScroll the transcript or rely on memoryTap the fact, read the utterance, replay the audio
Unsupported textHidden inside fluent sentencesHeld in a separate lane, out of the note
Correcting an errorEdit the note, then the chart, then the exportReject once; every surface follows
Evidence afterwardsThe note onlyThe note, its sources and the review decisions

The last row matters when questions come later, from a colleague, a patient or an insurer. A traceable note shows what was said and who approved each fact, not just what was written.

Fig. 02 · Comparison

Two ways to review a draft

Proofreading a note

  • Review a whole paragraph
  • Scroll the transcript for doubts
  • Unsupported text hides in prose
  • Fix note, chart and export separately

Reviewing traceable facts

  • Review one fact at a time
  • Tap, read the utterance, replay
  • Unsupported text held in a separate lane
  • Reject once; every surface follows
Traceability turns review from proofreading prose into checking single facts against their source.

Review habits that catch what AI misses

Design helps, but clinicians still need a routine. These habits take seconds and catch most high-risk errors:

  • Always check the high-risk facts: medications, doses, allergies, laterality, and anything you did not personally examine.
  • Read the negatives. Make sure the “denies” findings are things the patient actually denied.
  • Look for exam findings you did not do. If you did not examine it, it should not be in the note.
  • Tap to prove anything surprising. If a fact makes you pause, check the utterance before accepting it.
  • Report patterns, not just errors. If the same kind of mistake recurs in your room, tell the vendor and your colleagues. It is often the microphone or the room, not the model.
Fig. 03 · Checklist

Review habits that catch errors

  • Check medications, doses and allergies
  • Check laterality
  • Read every “denies” finding
  • Remove exams you did not do
  • Tap to prove anything surprising
  • Report recurring patterns
A few seconds of routine review catches most high-risk errors before you sign.

Where accountability sits

The signing clinician remains responsible for the note. Good governance makes that realistic: an audit trail of who accepted what, clear policies on what must be read, and a way to pause the AI if something goes wrong. Our guide to AI governance for clinics covers the wider controls. On privacy, including where audio and transcripts are stored, see the Helix security and data residency page.

How accurate are AI medical scribes?

It depends on the recordings, languages, specialty and how accuracy is measured. A single percentage says little about safety. Test any scribe on your own recordings and check whether you can trace every fact to its source.

What is an AI scribe hallucination?

It is a statement in the draft note that was never said in the visit, such as an exam finding that was not performed. Language models can produce these because they write what is typical for the presentation.

How does Verto stop invented facts reaching the note?

Every fact is linked to the utterance and audio it came from. Facts that cannot be traced are held in a separate lane and never reach the note unless a clinician confirms them.

Should I trust an AI scribe that quotes a high accuracy percentage?

Ask what was measured, on which recordings, and how errors were weighted. A high figure on clean English audio says little about a bilingual clinic. Test on your own visits and check whether errors are easy to catch.

Who is responsible for errors in an AI-drafted note?

The clinician who signs the note. The scribe is a drafting tool, which is why fast, evidence-based review matters.