RadMyk
← All posts
Guide September 13, 2026 · 7 min

Medical dictation security: who stores your radiology audio

Cloud medical dictation sends patient audio to vendor servers. Here's what gets stored, the breach risk, and how on-device dictation has no audio to expose.

By The RadMyk team

Every time a radiologist dictates into Dragon Medical One, that voice recording travels to a Microsoft Azure data centre, gets transcribed, and the text comes back. The round-trip is fast enough to feel instant. The medical dictation security question most radiologists never ask is what happens to the audio on the server side: how long it stays there, who can access it, and what it exposes if the vendor’s systems are ever compromised.

The compliance answer is well-rehearsed: cloud dictation vendors sign HIPAA Business Associate Agreements and maintain security programmes. The HIPAA compliance angle is covered separately. This post is about the security question underneath that. Even inside a compliant framework, cloud audio storage creates a breach surface that on-device dictation does not. The two concepts are not the same, and the distinction matters.

Does cloud medical dictation send audio to a server?

Yes. Every major cloud medical dictation product routes voice audio to a remote server. Dragon Medical One sends audio to Microsoft Azure. PowerScribe One uses the same Azure infrastructure. Augnito sends audio to Augnito’s cloud. Dolbey Fusion Narrate routes through nVoq, a third-party cloud speech engine under contract with Dolbey.

The speech recognition model and the language processing run on vendor infrastructure. Your device captures the audio and streams it out; the vendor’s server returns text. That is the architecture, and it works well when the network is reliable. It also means that patient audio leaves the reading room on every dictation, deposited on vendor hardware for at least as long as the vendor needs to process and retain it.

What do cloud vendors do with dictation audio after transcription?

The audio does not vanish the moment the text comes back. Cloud dictation vendors retain audio for several purposes.

Quality assurance. Most vendors retain audio samples to verify transcription accuracy, particularly when the system flags a low-confidence result. Vendor personnel may listen to the recording to assess model errors.

Model training. Under some data processing agreements, audio may be used, in anonymised or de-identified form, to improve the underlying speech model. Terms vary by contract, but it is a common practice in cloud speech services and worth reviewing in any BAA or data processing agreement before signing.

Audit and dispute resolution. Enterprise agreements often require vendors to retain audio so that the original dictation can be matched to the transcribed text if a dispute arises. The retention window in these cases can extend well beyond the processing window.

Sub-processor retention. When a cloud dictation vendor uses a third-party speech engine, the audio also travels to that sub-processor’s systems. Dolbey Fusion Narrate routes through nVoq. Nuance products route through Azure. The sub-processor’s retention policies are governed by the primary vendor’s BAA, but each party in the chain holds audio for some defined period.

None of this is hidden. Vendors describe these practices in their data processing documentation, and the BAA exists to govern them. The point is that patient audio containing protected health information does not disappear after transcription. It persists on vendor infrastructure, accessible to vendor personnel under controlled conditions, and stored on systems that are targets for external attack.

What does a breach expose in cloud radiology dictation?

If a cloud dictation vendor’s systems are compromised, the attacker’s access includes voice recordings of patient dictations.

Those recordings are not generic audio. A radiology dictation recording is specific: the patient’s study indication, the radiologist’s findings, the organ systems evaluated, any clinical history the radiologist includes, and the radiologist’s impression. Where the radiologist states or implies the patient’s name, study accession number, or date of birth, the recording becomes directly identifiable PHI tied to clinical findings, in voice format.

This is distinct from a breach of a text database. Transcribed text can sometimes be de-identified in processing. Voice recordings cannot. They carry the radiologist’s voice, prosody, and the full spoken content. They are also harder to process at scale: an attacker with a large set of audio recordings can extract information that might be obscured in a structured text export.

The scale of the exposure depends on the retention window. A vendor retaining 90 days of audio from a mid-sized radiology group holds a meaningful volume of patient dictations. A breach on day 89 exposes nearly three months of clinical audio. A vendor retaining audio for a shorter window, or not retaining it beyond the processing moment, presents a smaller target.

The sub-processor chain multiplies the exposure surface. If Dolbey Fusion Narrate routes audio to nVoq and nVoq is compromised, the dictation recordings of Dolbey customers who dictated during the retention window are in scope. The primary vendor’s BAA addresses legal liability, but it does not undo the disclosure.

Does HIPAA compliance make cloud dictation secure against breach?

HIPAA compliance reduces legal liability. It does not prevent a breach.

A vendor with a mature compliance programme, encryption in transit and at rest, and regular penetration testing can still be compromised. The healthcare sector has seen significant breaches at organisations with exactly those controls in place. Compliance is a legal floor, not a security guarantee.

The relevant question for a security assessment is not “is this vendor HIPAA-compliant” but “what does the breach surface look like, and how much does it contain?” HIPAA frameworks do not answer the second question. They answer the first.

For cloud medical dictation, the breach surface includes: voice recordings, transcribed text, associated metadata (timestamps, user identifiers, study accession numbers), and the sub-processor systems that handled the audio. Every radiologist dictating into a cloud platform is contributing patient audio to that surface for the duration of the retention window.

What changes when dictation runs on-device?

On-device dictation removes the cloud audio storage problem entirely. The speech model runs on the local machine. Audio is captured, processed on-device, and converted to text without leaving the device. No vendor server receives the recording.

If no audio leaves the machine, there is no audio on a vendor’s systems to retain, audit, or expose. There is no retention policy to evaluate because there is nothing to retain. The sub-processor chain does not exist for the dictation step because no third party is in the audio path.

This is not a reduction in the breach surface for the dictation layer. It is the elimination of it. Patient audio stays in the reading room because there is no mechanism for it to travel anywhere else.

The EHR, PACS, and RIS where the dictated text lands still carry their own security obligations. On-device dictation removes the dictation tool from that picture without affecting any other system. The offline dictation guide covers the related benefit: an on-device tool also keeps working when the network fails, which is a separate but related advantage.

How does RadMyk handle medical dictation security?

RadMyk processes audio locally on the radiologist’s machine. The speech recognition model runs on-device on macOS Apple Silicon and Windows. Voice never leaves the computer.

Because RadMyk’s servers never receive patient audio, there is no patient audio there to retain, to audit, or to expose in a breach. The vendor’s security posture is irrelevant to the dictation step because the vendor is not in that data path.

The on-device model is tuned for radiology vocabulary: anatomy, imaging modalities, measurement conventions, laterality, and structured report cadence. Out-of-the-box word accuracy is 96.1%, measured on radiology speech. Transcription runs at approximately 220 words per minute, keeping pace with full dictation throughput.

RadMyk types at the cursor in whatever application is active: PACS report fields, RIS text boxes, EHR text boxes, browser-based reporting tools, Microsoft Word, and Citrix or remote desktop sessions. No PACS integration project, no EHR plugin, no change to the reporting interface. If you can type into it, you can speak into it. The medical dictation software buyer’s guide covers how this cursor-based model compares to the integration-dependent enterprise platforms.

When is cloud dictation still the right choice?

For practices with reliable connectivity, existing enterprise agreements, and a compliance team already managing vendor BAAs, cloud dictation is a legitimate path. Dragon Medical One, Augnito, and PowerScribe One are used by thousands of radiology practices because the compliance frameworks work and the security programmes are real. The breach risk is not zero, but it is managed within a well-understood framework.

For practices where the breach surface itself is the concern, whether because of the sensitivity of the patient population, the volume of dictation audio that would be in scope, or institutional policies about outbound audio transmission, on-device dictation removes the question from the board.

The honest trade-off: cloud platforms offer portable profiles, central model management, and features like ambient AI and structured reporting templates that on-device tools do not match. On-device dictation offers a security architecture where the vendor never touches patient audio, because there is nothing to touch.

The bottom line on dictation audio security

Cloud medical dictation is built on an architecture that requires patient audio to reach a server. That architecture is HIPAA-compliant, vendor-managed, and widely deployed. It also means that patient voice recordings containing PHI exist on vendor infrastructure for as long as the retention window holds.

On-device dictation starts from a different place. The audio is processed on the device. The vendor’s servers do not receive it. There is nothing to breach in the dictation path because the vendor holds no audio.

RadMyk is one-time purchase, on-device dictation for radiologists. $199 one-time for the first 100 users (₹14,999 in India), rising after that. Radiology trainees use it free for the full length of their training. Practicing radiologists get a 28-day free trial, no credit card required.

See trial and pricing details.

VOICE

Own your voice. Pay once.

Start the free 28-day trial today, then own RadMyk with a single payment. No subscription, ever. And if you're still training, it's yours for free.