How to Transcribe Audio in Microsoft Word: Full Guide
You have an interview, a lecture recording, or a voice memo sitting on your desktop, and typing it out by hand would take hours. Microsoft Word has a built-in Transcribe feature that turns an audio file into timestamped, speaker-separated text without leaving the document. It is genuinely useful once you know where it lives and what it refuses to do.
This guide walks through how to transcribe audio in Microsoft Word using the native tool, then covers the limits you will hit and what to do about them. Everything below reflects the current Word for the web and Microsoft 365 desktop behaviour.
What You Need Before You Start
The Transcribe feature is not part of every Word installation. Three things have to be true:
- You have a Microsoft 365 subscription that includes the Transcribe feature. It is not available in one-time-purchase versions of Word such as Word 2019 or Word 2021.
- You are working in Word for the web, or in the Word desktop app on Windows or Mac while signed in to your Microsoft 365 account. Transcribe runs in the cloud, so the file uploads to OneDrive either way.
- Your browser is Edge or Chrome if you are using Word for the web. Other browsers may not expose the Transcribe button at all.
Supported audio formats are MP3, WAV, M4A, and MP4. If your file is in another format, convert it first or you will not get past the file picker.
Step 1: Open the Transcribe Pane in Word
Open your document, or create a blank one to hold the transcript. Go to the Home tab on the ribbon. On the right-hand side you will find a Dictate button with a small arrow beneath it.
Click that arrow to open the dropdown, then select Transcribe. A pane opens on the right side of the window with two options: Upload audio and Start recording.
If you do not see Dictate at all, you are either not signed in or your licence does not include the feature. Sign out and back in to your Microsoft 365 account, then reload the page before assuming it is missing.
Step 2: Choose Your Transcription Language
Before uploading anything, look at the top of the Transcribe pane. There is a language selector, usually defaulted to the language of your Office installation.
Click it and pick the spoken language of your recording. This is the single most common cause of unusable output: an English-language interface transcribing a Spanish recording will produce nonsense that looks like real words. Word does not auto-detect language, and it will not warn you that something is wrong.
Word supports a reasonable range of languages here, though the list is shorter than the list of languages Office itself ships in. If your language is not present, the native route will not work and you will need an external tool.
Step 3: Upload Your Audio File
Click Upload audio. A standard file picker opens. Select your MP3, WAV, M4A, or MP4 file and confirm.
Word uploads the file to a folder in your OneDrive called Transcribed Files. This matters for two reasons. First, the audio consumes your OneDrive storage quota, so a long recording will eat into it. Second, if you later delete that OneDrive file, the transcript in your document keeps its text but loses playback, so the timestamps stop being clickable.
Transcription now runs in the cloud. Roughly speaking, expect to wait about as long as the recording itself, sometimes less. You must keep the tab open while it processes. Closing the browser or putting the machine to sleep will cancel the job, and you will have to start again from the upload step.
Step 4: Review the Transcript and Fix Speaker Labels
When processing finishes, the pane fills with the transcript, broken into sections. Each section carries a timestamp, a speaker label such as Speaker 1, and the transcribed text.
Click any timestamp to jump the audio player to that point. This is the fastest way to verify a passage you suspect is wrong: read, click, listen, correct.
To fix errors, hover over a section and click the pencil icon. You can edit the text directly. To rename a speaker, click the pencil next to the speaker label; Word offers to apply the new name to every instance of that speaker in the transcript, which saves a lot of repetitive clicking.
Word separates speakers reasonably well when people take turns cleanly, and poorly when they talk over each other. Overlapping speech usually gets merged into one speaker's block, and you will need to split it manually.
Step 5: Insert the Transcript Into Your Document
The transcript lives in the side pane, not the document, until you put it there. At the bottom of the pane, click Add to document. You get four choices:
- Just text gives you clean prose with no timestamps or speaker names. Best for turning an interview into an article.
- With speakers adds the speaker labels but drops the timestamps. Good for scripts and dialogue.
- With timestamps keeps the time markers but not the names. Useful for finding moments in the original recording.
- With speakers and timestamps gives you everything, which is what you want for meeting minutes or legal-style notes.
Pick one and the text drops into your document at the cursor. From there it is ordinary Word text: styled, searchable, and exportable.
Step 6: Handle Long or Multiple Recordings
Word caps a single uploaded file at 300 MB, and Microsoft applies a monthly limit of 300 minutes of uploaded transcription per account. Both limits are easy to hit if you transcribe regularly.
For a file over 300 MB, split the audio into parts using any audio editor, transcribe each part separately, and paste the results into one document in order. Remember that timestamps restart at zero for each part, so note the offset for each segment before you merge.
If you are regularly over the monthly minutes, or you need output as SRT or VTT subtitle files rather than Word text, the native feature stops being the right tool. This is where a dedicated transcription service earns its place. Tapescribe takes a file or a link, returns the transcript plus SRT, VTT, and TXT exports, and adds chapter markers, which Word does not produce at all. For a broader look at the options, our guide to transcribing audio to text online compares the general approaches.
Alternative: Dictate Live Instead of Uploading
If the audio does not exist yet, skip the upload path entirely. In the same Transcribe pane, click Start recording and Word transcribes as you speak, saving the audio to OneDrive when you stop.
There is also plain Dictate, the button itself rather than its dropdown. Dictate types directly into your document in real time with no audio file and no transcript pane. It does not separate speakers and does not save a recording, so it suits drafting rather than documenting a conversation.
One practical trick: recording your own voice re-reading a difficult passage often produces cleaner text than transcribing the poor-quality original.
Common Problems and How to Fix Them
The Transcribe option is greyed out or missing. Almost always a licensing issue. Confirm you are on a Microsoft 365 subscription rather than a perpetual licence, and confirm you are signed in with that account and not a personal one.
Transcription stalls or never completes. The tab was backgrounded, the machine slept, or the connection dropped. Keep the tab in focus, disable sleep for long files, and retry. Very long files fail more often than short ones, so splitting is a reliable workaround.
The output is gibberish. Check the language selector first. If the language is right, the problem is audio quality: background noise, distant microphones, and heavy compression all degrade accuracy sharply.
Speakers are merged or mislabelled. Word struggles with crosstalk and with more than three or four voices. Recording each participant on a separate track and transcribing them individually is the only real fix. This comes up constantly with meeting recordings; if that is your use case, the Microsoft Teams transcription guide covers the Teams-native route, which handles multi-speaker calls better because it knows who is on the call.
You need subtitles, not a document. Word has no subtitle export. You would have to rebuild timing by hand, which is not worth doing. Use a tool that outputs SRT or VTT directly. The same applies if you need timestamps for a video rather than a written record.
The file format is rejected. Only MP3, WAV, M4A, and MP4 work. Convert anything else, and note that a video file in MOV or MKV will need converting even though MP4 video is accepted.
Frequently Asked Questions
Is transcribing audio in Microsoft Word free?
Not on its own. Transcribe requires a Microsoft 365 subscription, so it is included in a plan you already pay for rather than being a free standalone tool. Perpetual licences such as Word 2021 do not include it, and there is no way to add it separately.
How long can an audio file be to transcribe in Word?
A single file can be up to 300 MB, and your account gets 300 minutes of uploaded transcription per month. Duration is limited by file size rather than a stated time cap, so a heavily compressed MP3 can run considerably longer than a WAV of the same size. Split larger recordings into parts and transcribe them one at a time.
Can Microsoft Word transcribe audio without an internet connection?
No. Both Transcribe and Dictate send audio to Microsoft's cloud services for processing, so an active connection is required throughout. Losing connectivity mid-job cancels it, and you will need to upload the file again from the start.
Does Word transcription identify different speakers automatically?
Yes, Word attempts speaker separation and labels each one as Speaker 1, Speaker 2, and so on. Accuracy depends heavily on how cleanly people take turns; overlapping speech tends to get attributed to a single speaker. You can rename each label once and Word will apply that name throughout the transcript.
Practical Takeaway
Word's Transcribe feature is the right choice when you already pay for Microsoft 365, your recording is clean, and your end product is a document. Set the language before you upload, keep the tab open while it processes, and use the timestamp playback to verify anything that reads oddly. When you outgrow the monthly minutes, need subtitle files, or are working with messy multi-speaker audio, that is the signal to move the job somewhere purpose-built rather than fighting the limits.