Back to blog

How to Transcribe iPhone Voice Memos to Text: Full Guide

TranscriptioniPhoneWorkflow

You recorded something important in Voice Memos, an interview, a lecture, a podcast segment, a two hour meeting, and now you need it as text you can search, edit, or quote. The iPhone can do some of this on its own, but the built-in feature has real limits that bite exactly when the recording matters most. This guide covers how to transcribe iPhone Voice Memos to text three ways: the native iOS transcript, exporting the audio file to a proper transcription tool, and getting timestamped or speaker separated output for editing.

Start with Step 1 if you are on a recent iPhone and the recording is short. Skip to Step 3 if your recording is long, has multiple speakers, or you need SRT or timestamps.

Step 1: Use the built-in Voice Memos transcript in iOS 18

Apple added automatic transcription to Voice Memos in iOS 18. It runs on device, it is free, and it needs no setup beyond a supported iPhone and a supported language.

  1. Open the Voice Memos app.
  2. Tap the recording you want in the list so it expands.
  3. Tap the three dots button (More) at the bottom right of the expanded player.
  4. Choose View Transcript. On some builds the transcript icon appears directly in the player as a small speech bubble with lines.
  5. Wait a few seconds. For a short memo the text appears almost immediately; for a longer one you will see it fill in progressively.

To copy the text out:

  1. With the transcript open, tap the three dots again.
  2. Tap Copy Transcript.
  3. Paste into Notes, Mail, Slack, or wherever you need it.

If View Transcript is missing entirely, check three things. First, your iPhone must be running iOS 18 or later and be a model that supports the feature; older devices do not get it even on the right iOS version. Second, the recording language must be one Apple supports for transcription, and it must roughly match your Siri or device language setting. Third, transcription is generated per recording, so memos captured before you updated may need you to open them once and wait while the system processes them.

Step 2: Check whether the native transcript is actually good enough

Before you build a workflow on the built-in transcript, read the output critically. The native feature is genuinely handy for quick voice notes, but it has known weak points:

  • No speaker labels. Everything comes out as one continuous block of text. For an interview or a two person meeting, you cannot tell who said what without going back to the audio.
  • No timestamps you can export. The app highlights words as playback moves, but copying the transcript gives you plain prose with no time codes attached.
  • No subtitle formats. There is no SRT or VTT export, so if the audio is destined for a video, you are converting by hand.
  • Weak with accents, crosstalk, and background noise. On-device models are small by design. Cafe recordings, phone speaker audio, and heavy accents degrade quickly.
  • It struggles with long recordings. More on that below.

If the memo is you dictating a shopping list or a two minute idea, Step 1 is the whole answer and you can stop reading. If it is source material for work, keep going.

Step 3: Export the Voice Memo as an M4A file

To use any external transcription tool, you need the audio file itself. Voice Memos records to M4A, which is a standard, widely supported format. There are three reliable ways to get it off the phone.

Option A: Share sheet to Files (best for uploading later)

  1. In Voice Memos, tap the recording, then tap the three dots.
  2. Tap Share.
  3. Choose Save to Files.
  4. Pick a location, iCloud Drive or On My iPhone, and tap Save.

The file lands as a .m4a you can upload from the Files app in any browser or app.

Option B: AirDrop to a Mac (best for editing on a computer)

  1. Tap the recording, then three dots, then Share.
  2. Tap the AirDrop icon and select your Mac.
  3. On the Mac, accept if prompted. The file arrives in Downloads.

If AirDrop does not find the Mac, make sure both devices have Wi-Fi and Bluetooth on, and that the Mac's AirDrop receiving setting is not restricted to Contacts Only when the devices use different Apple accounts.

Option C: Share to email or cloud storage

Same share sheet, but choose Mail, Google Drive, Dropbox, or your messaging app of choice. Note that email providers cap attachment sizes, and a long uncompressed voice memo can exceed that. Cloud storage or Save to Files is safer for anything over roughly an hour.

One thing worth knowing: you do not need to convert the M4A to MP3 or WAV first. Most transcription services accept M4A directly, and re-encoding only costs you quality and time. If a tool rejects it, that is a limitation of that tool, not of your file.

Step 4: Transcribe the file with a tool that handles length and speakers

Once you have the M4A, upload it to a transcription service. This is where you solve the problems Step 1 could not. If you regularly work with recorded audio from several sources, the same approach applies to any file; the general workflow for transcribing audio to text online is identical whether the source is a voice memo, a call recording, or a video's audio track.

What to look for, in priority order:

  1. Length tolerance. The tool should accept a full length recording without asking you to split it. Ask this before you upload two hours of audio.
  2. Timestamps. You want time codes at least at the paragraph level, ideally per segment, so you can jump back to any claim in the audio.
  3. Speaker separation. For interviews, the transcript should mark speaker changes rather than running everything together.
  4. Export formats. TXT for reading and quoting, SRT or VTT if the audio is going into a video.

Tapescribe takes an M4A upload directly and returns a timestamped transcript along with SRT, VTT, and TXT exports, which covers the formats you would otherwise have to build by hand. It has a free tier, so you can run one memo through and compare the output against your iOS transcript before deciding whether it is worth using regularly.

Step 5: Get timestamped output you can actually edit against

If the memo is reference audio for a video edit, a podcast, or a written piece, plain prose is not enough. You want to be able to read a line and know exactly where it sits in the recording.

For a written piece or notes: export TXT with timestamps. Paste it into your document, then delete the timestamps for lines you are not going to use. What is left is a set of quotes each tied to a time you can verify.

For a video or podcast edit: export SRT. Most editors, including Premiere Pro, DaVinci Resolve, Final Cut Pro, and CapCut, will import an SRT as a caption or subtitle track. Even if you never publish those captions, having the text laid out on the timeline makes finding a specific sentence far faster than scrubbing. If your final destination is short-form video, the same SRT is the basis for burned-in captions; the process is the same one covered in the guide to transcribing TikTok videos to text.

For a searchable archive: export TXT, name the file to match the recording date and topic, and keep transcripts in one folder. Text search across a folder of transcripts is dramatically faster than trying to remember which memo contained what.

Step 6: Handle interviews with two or more speakers

A recorded interview is the case where the native transcript fails hardest, because a single undifferentiated block of text is close to unusable for quoting.

Three things improve the result before you even transcribe:

  1. Place the phone between speakers, not next to one. A memo dominated by one voice makes separation harder for any tool.
  2. Have each person say their name at the start. This gives you an anchor to map speaker labels to real names in the transcript.
  3. Avoid talking over each other where you can. Crosstalk is the single biggest source of errors in multi-speaker audio.

After transcribing, do a cleanup pass: find and replace the generic speaker labels with real names, then skim for the sections where speakers overlap and fix those by ear. For a fuller treatment of the whole process from recording to a clean quotable document, see the guide on how to transcribe an interview to text.

Common problems and how to fix them

The transcript option does not appear. Your iPhone model or iOS version does not support native transcription, or the recording language is unsupported. Export the M4A and use an external tool instead.

Long recordings never finish transcribing. On-device transcription is memory and battery constrained, and very long recordings can stall, time out, or produce a partial transcript. Keep the phone plugged in, keep Voice Memos in the foreground, and give it time. If it still fails, export the file; server side transcription is not subject to the same constraint.

The transcript is full of errors. Check the recording first. Phone microphone audio, distance from the speaker, background noise, and heavy accents all degrade accuracy. A cleaner recording beats any post-processing fix.

The M4A is too large to email. Use Save to Files and upload from there, or share via cloud storage. Do not compress the audio to make it fit; you will lose accuracy.

Names, jargon, and acronyms come out wrong. No transcription system knows your company's product names. Build a find and replace list of your recurring terms and run it over every transcript. It takes two minutes and fixes the same errors permanently.

You need the transcript from a voice note that was not recorded in Voice Memos. Messaging app audio is a different pipeline; the approach for WhatsApp voice notes covers exporting from the app first.

Frequently Asked Questions

Can iPhone transcribe voice memos automatically without an app?

Yes, on iOS 18 and later on supported iPhone models. Open the recording in Voice Memos, tap the three dots, and choose View Transcript. The transcription happens on device at no cost, but it produces plain text only, with no speaker labels, no exportable timestamps, and no subtitle file formats.

Why does my long voice memo fail to transcribe on iPhone?

On-device transcription runs within tight memory and battery limits, so a recording that runs for an hour or more can stall or return only part of the text. Keeping the phone charged and the app open helps, but the constraint is real. Exporting the M4A and transcribing it with a server side tool avoids the problem entirely, because the processing does not happen on your phone.

What file format do iPhone Voice Memos use?

Voice Memos saves recordings as M4A, an AAC audio container that is widely supported. You do not need to convert it before transcribing; most transcription tools accept M4A as is. Converting to MP3 or WAV first adds a step and can slightly degrade the audio, which works against accuracy.

How do I get a transcript with speaker names from an interview recording?

The native iOS transcript does not separate speakers, so you need a tool that supports speaker detection. Export the M4A from Voice Memos, upload it to a transcription service that labels speaker changes, then find and replace the generic labels with real names. Recording with the phone placed between speakers and having each person state their name at the start makes the labels far easier to map correctly.

Pick the method that matches the stakes of the recording. A quick note to yourself needs nothing more than the built-in transcript and a copy paste; an interview, a client call, or source audio for a video is worth exporting as M4A and running through a tool that gives you timestamps, speaker separation, and an SRT you can drop straight onto a timeline. Either way, the first move is the same: get the file off the phone before you need it, because the worst time to discover a transcription limit is the day the deadline lands.