Radiology dictation accuracy: what the numbers mean
Vendor accuracy claims for medical dictation software are rarely measured the same way. Here is what to look for and why measurement conditions matter in radiology.
By The RadMyk team
When you evaluate radiology dictation software, every vendor hands you an accuracy figure. Some quote a word error rate. Some claim recognition above 99%. Most do not say how the number was measured, in what conditions, or at what stage of the product’s adaptation to a particular speaker.
The difference matters. Radiology voice recognition software that achieves 99% in a quiet room with a professional microphone after six months of profile training will perform differently on day one in a busy reading room with a basic USB headset.
This guide explains what accuracy numbers mean for medical dictation software, how cloud and on-device approaches differ in practice, and what to look for when a vendor quotes you a percentage.
What does word accuracy mean for dictation software?
Word error rate (WER) is the standard measure: the share of words the software gets wrong when transcribing a given input. A WER of 3.9% means 96.1% of words are correct. A WER of 7% means accuracy is 93%.
That sounds precise. Across vendors it is not directly comparable, because three variables shift the result substantially:
The test vocabulary. Accuracy measured on a general medical word set is not the same as accuracy on radiology-specific text. “Perifissural nodule,” “no acute cardiopulmonary process,” and “osseous structures are intact” stress a recognition model differently than standard clinical phrases. A vendor that measures accuracy on general clinical dictation may score well while struggling on radiology subspecialty terminology.
The measurement conditions. A controlled recording session with a calibrated microphone, no background noise, and a single rested speaker is not a reading room. Reading rooms have HVAC noise, nearby conversations, phone alerts, and radiologists who dictate while scrolling through a PACS worklist. If a vendor does not specify the recording environment, the number is aspirational.
The adaptation state. Cloud dictation tools often improve as they accumulate data about a specific speaker’s voice, vocabulary preferences, and microphone characteristics. A figure achieved after six months of daily use is not what a new user gets on day one.
The only accuracy number that means anything to a radiologist is what the software produces, on radiology vocabulary, in real reading room conditions, from the first day of use. Most vendor claims do not specify all three conditions.
Does cloud processing improve dictation accuracy?
Cloud dictation tools (Dragon Medical One, Augnito, Dolbey Fusion Narrate, PowerScribe One) process audio on vendor servers rather than on the local machine. Those servers run larger speech models than a personal workstation could support, and the model parameters can be updated centrally without requiring a software re-download.
This is a genuine technical advantage for cloud vendors. Larger models handle unusual vocabulary and accented speech more consistently. Adaptive profiles, stored in the cloud, follow the radiologist across devices, which matters for those who read across multiple hospital systems.
Dragon Medical One demonstrates this well. Its adaptive profile improves steadily with use, and the model has been trained on a large clinical corpus across specialties. Its KLAS recognition record reflects real-world deployment at scale. Augnito maintains specialized models for over 55 clinical specialties, including radiology, and updates those models without requiring action from the user.
These are honest strengths. A well-adapted cloud profile, after months of use, can reach accuracy that a new on-device installation would need calibration to match.
The limitation appears at the edges. Cloud accuracy depends on the audio reaching the server and the text returning. When the network is slow, the round-trip becomes noticeable. When the network drops, the tool stops. A radiologist in a reading room with unreliable VPN connectivity does not get a degraded experience from cloud dictation - they get no experience at all. The failure mode is binary, not graceful.
Cloud tools also require an adaptation or enrollment period before reaching their best performance. The vendor quotes the post-adaptation figure. The first week tends to be rougher than the marketing suggests.
How does on-device processing handle accuracy?
On-device speech recognition runs the model locally. No audio leaves the machine. The model is loaded once during setup and processes each dictation without a network round-trip.
The accuracy ceiling is, in theory, lower than a data-centre deployment, because the model must fit on consumer hardware. The floor, however, is more consistent. On-device accuracy does not degrade when the VPN drops, does not vary with server load at the vendor’s data centre, and does not require a multi-month adaptation window before it becomes reliable.
RadMyk’s current speech model achieves 96.1% word accuracy on radiology vocabulary, measured out of the box, before any calibration, in representative reading room conditions. This figure comes from RadMyk’s own benchmark, not a vendor-controlled test, and it is the accuracy a new user experiences on the first day.
An optional guided calibration step, which takes one short session and tunes the model to a specific voice and microphone, improves accuracy further. But the out-of-the-box number is the honest starting figure, not the aspirational one after ideal conditions and months of use.
Does on-device dictation require a long training period before it is usable?
Older on-device recognition systems shipped with generic acoustic models and required hours of recorded training material before becoming accurate enough for clinical use. That reputation has followed on-device dictation long past its expiry date.
Modern radiology-tuned on-device tools ship pre-trained models. RadMyk’s model is trained on radiology language before it reaches the radiologist: anatomy, imaging modalities, measurement conventions, laterality terms, and the cadence of structured reporting. A radiologist does not need to read aloud for two hours before the software understands “right lower lobe consolidation.”
The guided calibration step in RadMyk is a short recording session, not a multi-hour enrollment. After calibration, no further training is required unless the microphone setup changes. The model’s vocabulary does not need teaching; the calibration is voice and hardware tuning, not curriculum.
Cloud tools, by contrast, adapt continuously by storing evolving voice profiles on vendor servers. That ongoing adaptation is part of the value proposition, but it also means the profile lives outside the radiologist’s control. When a vendor migrates platforms - as happened when PowerScribe 360 announced its end-of-life and users began moving to PowerScribe One - the question of what happens to existing voice profiles, macros, and customizations is not always answered clearly in advance.
With on-device processing, the model is on the radiologist’s machine. No migration. No profile transfer. No re-enrollment if the vendor changes architecture.
What does 96.1% accuracy mean across a full reading list?
For a radiologist dictating 30 reports in a session, at roughly 150 to 200 words per report, that is 4,500 to 6,000 words per day. At 96.1% accuracy, roughly 175 to 234 words across the session need correction. At 93% accuracy, the correction count rises to 315 to 420 words.
Over a full reading week, that gap compounds. Correction time is not only the time to fix the word. It is the pause, the cursor movement, the reread, and the re-entry into dictation flow. Reported transcription error rates appear in multiple studies as a driver of the documentation burden that contributes to radiologist burnout.
RadMyk also transcribes at approximately 220 words per minute, which matches full dictation speed without constraining the radiologist to speak slowly to maintain recognition quality.
Which tools get accuracy right for radiology?
The honest answer is that several tools achieve acceptable radiology accuracy. The right choice depends on more than the number.
Dragon Medical One’s adaptive cloud profile is strong after the enrollment period, and its KLAS recognition reflects genuine clinical deployment experience. The absence of a native Mac client is a real constraint for radiologists on Apple hardware.
Augnito has radiology-tuned models and a strong presence in international markets. It is a cloud product: audio travels to Augnito’s servers, and the subscription model runs per user per year.
M*Modal Fluency for Imaging offers CAPD nudges, structured reporting, and actionable findings management. It is an enterprise platform with a scope that goes well beyond front-end dictation. If those capabilities are needed, RadMyk does not offer them.
An honest comparison across the main dictation options for radiology is worth reading before committing to any product in this category.
What to measure during a trial
RadMyk is available at radmyk.com with a 28-day free trial that requires no credit card and no email verification. Radiology trainees (residents and fellows) use RadMyk free for the entire length of their training, with no countdown.
For practicing radiologists, the trial runs long enough to measure accuracy across a real reporting schedule: several modalities, different study types, early-morning and end-of-shift dictation. The number to track is not what the vendor claimed. It is how many corrections you make per report, whether that decreases over the first week, and whether the software keeps working when the network does not.
For those who want to buy rather than trial, RadMyk is $199 one-time for the first 100 users (Rs 14,999 in India for the first 100 users), with the list price rising after that offer closes. One payment. No renewal. No subscription next year. The full pricing details are on the pricing page.
Accuracy is the right question to ask. Make sure you are measuring the number that applies to your reading room, not the vendor’s controlled test.