RadMyk

/ benchmark

How accurately does it transcribe radiology?

An open benchmark for radiology speech-to-text. 100 radiology report sentences are spoken by 10 AI voices across 5 accents and degraded into 4 audio conditions, then transcribed by RadMyk's on-device model and scored against the exact text. Play any clip - words highlight as spoken - and see what the model heard.

3.9%

Word error rate (WER)

Share of words transcribed wrong vs. the exact sentence. Lower is better.

97%

Term accuracy

Share of key radiology terms caught (e.g. air-space consolidation). Higher is better.

Average across 10 voices and 4 audio conditions · model: RadMyk's radiology fine-tune of google/medasr, no per-user calibration · 400 clips

Results by audio condition

The same sentences and voices, played through progressively harder audio. Accuracy holds up on clean and phone-line audio and dips most under heavy reverb.

Condition What it simulates WER Term accuracy
Clean studio-quality audio 3.3% 100%
Noisy room background room noise 4.2% 97%
Reverb a reverberant room 4.8% 91%
Phone line compressed phone-line audio 3.2% 99%

Hear it for yourself

Pick a speaker and an audio condition, then play any clip. The reference words highlight as they're spoken; below, the model's transcript shows wrong and extra words.

Speaker
Audio
WER 3.0% term accuracy 100%
matched wrong word missed extra
Chest radiography and thoracic CT #0001

Reference

Portable chest radiograph shows a right upper lobe air-space consolidation, with air bronchograms and no pleural effusion.

Model heard

portable chest radiograph shows a right upper lobe air space airspace consolidation with air bronchograms and no pleural effusion

Chest radiography and thoracic CT #0002

Reference

Patchy bilateral lower lobe ground-glass opacities are present, greater on the left.

Model heard

patchy bilateral lower lobe ground glass opacities are present greater on the left

Chest radiography and thoracic CT #0003

Reference

There is diffuse smooth interlobular septal thickening consistent with pulmonary edema.

Model heard

there is diffuse smooth interlobular septal thickening consistent with pulmonary edema

Chest radiography and thoracic CT #0004

Reference

A small left pneumothorax is visible at the apex, without mediastinal shift.

Model heard

a small left pneumothorax is visible at the apex without mediastinal shift

Chest radiography and thoracic CT #0005

Reference

Moderate right pleural effusion causes adjacent compressive atelectatic change.

Model heard

moderate right pleural effusion causes adjacent compressive atelectatic change

Chest radiography and thoracic CT #0006

Reference

CT demonstrates lower lobe predominant subpleural honeycombing with traction bronchiectasis.

Model heard

ct demonstrates lower lobe predominant subpleural honeycombing with traction bronchiectasis

Chest radiography and thoracic CT #0007

Reference

There are scattered centrilobular tree-in-bud nodules in the right middle lobe and lingula.

Model heard

there are scattered centrilobular tree in bud nodules in the right middle lobe and lingula

Chest radiography and thoracic CT #0008

Reference

Thin-section CT shows diffuse bilateral mosaic attenuation that becomes more conspicuous on expiratory images.

Model heard

thin section ct shows diffuse bilateral mosaic attenuation that becomes more conspicuous on expiratory images

Chest radiography and thoracic CT #0009

Reference

A 7 mm solid pulmonary nodule is seen in the posterior segment of the right upper lobe.

Model heard

a 7 mm solid pulmonary nodule is seen in the posterior segment of the right upper lobe

Chest radiography and thoracic CT #0010

Reference

The left lower lobe contains a 2.1 cm cavitary lesion with a thick irregular wall.

Model heard

the left lower lobe contains a 2.1 cm cavitary lesion with a thick irregular wall

Don't take a chart's word for it.

Dictate a real radiology passage with your own engine and get your accuracy and words-per-minute scored live in the browser.

Test your dictation →

How the benchmark works

Each test sentence is a realistic radiology report line covering chest, cardiac, neuro, spine, and musculoskeletal imaging, with a marked key term that must be transcribed correctly. Every sentence is spoken by 10 AI voices (American, British, Australian, Canadian, Indian accents, both genders) so accuracy isn't measured on a single voice.

Each clip is then rendered in 4 conditions - clean, noisy room, reverb, phone line - and transcribed by the on-device model. We score word error rate against the exact sentence and term accuracy against the marked key phrase. Because the audio is synthetic and fixed, the score is fully reproducible: it moves only when the model changes.

Part two: real radiologists (in the works)

The synthetic benchmark above is deliberately fixed and reproducible - it moves only when the model moves. It is also, deliberately, not the whole evidence base. The next layer is a real-radiologist pilot, measured the way you actually work:

  • Practicing radiologists and residents, Indian and international accents
  • Real microphones: headsets, desktop mics, laptop mics
  • Real rooms: quiet reporting rooms and noisy shared ones
  • Real report passages - measurements, laterality, structured phrasing
  • Scored on word error rate, critical-term accuracy, number and measurement accuracy, and correction time

Results will publish on this page as they land. Until then, treat part one as the reproducible baseline it is - and run your own dictation through the test for the only sample that really matters: yours.

Questions

What is word error rate (WER)?

+

Word error rate is the percentage of words the speech-to-text model gets wrong - substitutions, deletions, and insertions - measured against the exact reference sentence. Lower is better; 0% is a perfect transcript. Across this benchmark the model averages 3.9% WER.

What is term accuracy?

+

Term accuracy is the share of clinically important radiology terms the model transcribes correctly - the key phrase in each sentence, such as “air-space consolidation” or “BI-RADS 4”. It measures what matters for a report, not every filler word, and treats hyphenation variants (“air-space” vs “airspace”) as correct.

How was this benchmark built?

+

100 radiology report sentences are each spoken by 10 AI voices spanning 5 accents and both genders, then degraded into 4 audio conditions. Every clip is transcribed by the model and scored against the exact reference sentence (400 clips in total).

Why synthetic AI voices instead of real recordings?

+

Synthetic audio is fixed and reproducible, so a change in the score reflects a change in the model rather than the recording, and it contains no patient data. The trade-off is that it is a controlled proxy, not a substitute for real-world dictation.

What audio conditions are tested?

+

Clean studio audio plus a noisy room, a reverberant room, and a compressed phone line - applied to every clip, so you can see how accuracy holds up as conditions degrade.

Which model produced these results?

+

These results are from the model that ships with RadMyk: a radiology fine-tune built on top of the google/medasr base. No per-user personalization or voice calibration was applied - this is out-of-the-box accuracy, as installed.

Does this reflect real-world dictation accuracy?

+

It is a strong indicator of how the model handles radiology vocabulary across voices and audio conditions, but real dictation adds live microphones, spontaneous speech, and individual accents. Treat it as a controlled benchmark, not a guarantee.

Can I download the audio or dataset?

+

No. The clips are a demo you can play on this page; they are not distributed as a downloadable dataset.

Synthetic benchmark · model: RadMyk's radiology fine-tune of google/medasr · 400 clips across 4 audio conditions. Voices are AI-generated for testing; this is a showcase, not a downloadable dataset.

VOICE

Own your voice. Pay once.

Start the free 28-day trial today, then own RadMyk with a single payment. No subscription, ever. And if you're still training, it's yours for free.