How to Transcribe Interviews for Qualitative Research (Full Guide)
You have a folder of interview recordings, a coding deadline, and a transcript standard your supervisor or ethics reviewer will actually check. Working out how to transcribe interviews for qualitative research is less about typing speed than about decisions you make before the first word hits the page: verbatim style, notation, timestamp granularity, pseudonyms. Get those right and your transcripts drop cleanly into NVivo, ATLAS.ti or Dovetail. Get them wrong and you will be recoding the same data twice.
This guide walks through the manual method first, then where automation genuinely saves time without compromising your data.
Step 1: Sort consent, storage and file naming before you transcribe
Open your ethics approval and your participant information sheet and check three things. First, whether transcription by a third party is permitted, and whether you promised participants that only the research team would hear the audio. If you did, an external service or a hired transcriber is off the table until you amend the approval. Second, where recordings may be stored, which usually means institutional storage rather than a personal cloud drive. Third, what you told participants about retention.
Then set up a participant ID scheme and use it everywhere: P01, P02, and so on. Rename recordings to something like P07_2026-03-04_int1.wav. Keep the key that maps IDs to real names in a separate encrypted file, not in the same folder as the audio, and never inside the transcript itself.
Step 2: Decide between verbatim and intelligent verbatim
This is the decision that shapes everything else, so make it once and write it into your methods section.
Full verbatim captures every utterance exactly: false starts, repetitions, "um" and "er", stutters, laughter, pauses, overlapping speech. Choose it if your analysis attends to how something is said. Conversation analysis, discourse analysis and interactional work all need it, usually with a formal system such as Jefferson notation, which marks pause length, latching, intonation and volume.
Intelligent verbatim, sometimes called clean verbatim, keeps every substantive word but removes filler sounds, stammers and obvious slips of the tongue. Nothing is paraphrased and no sentence is reordered. This is the standard choice for reflexive thematic analysis, framework analysis, grounded theory and most applied qualitative work, because it produces a readable transcript without discarding meaning.
Denaturalised transcription goes a step further and tidies grammar and dialect into standard written form. Use it sparingly; it can strip out exactly the voice you are trying to represent, and it raises fair questions about whose language you are privileging.
If your interviews are commercial rather than academic, the trade-offs differ and the standards are looser, and a lighter workflow like the one in our guide to transcribing customer interviews will serve you better than a full research protocol.
Step 3: Write your notation conventions into a template
Create one Word or Google Doc template and reuse it for every interview. At the top, record the interview ID, date, duration, setting, and the transcription convention you are using. Then fix your symbols and stick to them:
I:for interviewer,P07:for participant. Use the same labels in every file.[laughs],[sighs],[phone rings]for non-verbal sounds in square brackets.(.)for a short pause,(3)for a timed pause of three seconds.[inaudible 00:14:22]with a timestamp so you can go back and try again.[unclear: research assistant?]when you have a guess but are not confident.[city],[employer],[colleague name]as placeholders for identifying details.- Round brackets for your own clarifying insertions, so a reader can always tell what came from the participant.
Consistency matters more than which convention you pick, because your CAQDAS software will let you search on these strings later. A search for inaudible should return every gap in the corpus.
Step 4: Transcribe manually with a free player
The baseline method needs no paid tools. Open the audio in a player with keyboard control and type into your template beside it.
VLC works well. Use the [ and ] keys to slow playback down and speed it back up, Shift plus left arrow to jump back three seconds, and Alt plus left arrow for ten. Most people transcribe comfortably at somewhere between two thirds and three quarters of normal speed. A browser tool such as oTranscribe keeps the player and the text box in one window, so Esc pauses and resumes without your hands leaving the keyboard, and it can insert a timestamp with a shortcut. If you are transcribing regularly, a USB foot pedal with Express Scribe pays for itself in wrist strain avoided.
Work in passes rather than trying to be perfect in one go. First pass, get the words down and mark gaps. Second pass, listen again at normal speed with the transcript in front of you and fix errors, add non-verbal notation, and resolve the inaudible markers. Expect manual transcription to take several times longer than the recording itself, which is why most researchers now use step 5 for the first pass.
Step 5: Use automatic transcription for the first pass, then correct it
Automatic speech recognition is now accurate enough to produce a solid draft, and correcting a draft is far faster than typing from silence. The workflow is simple: upload the recording, wait for the draft, then play the audio back at normal speed while reading along and fixing errors.
Tapescribe handles this part, taking an uploaded file or a link and returning a timestamped transcript with speaker labels plus TXT, SRT and VTT exports, which matters at step 8. Whatever tool you use, treat the output as a draft, never as data. Automatic systems mishear proper nouns, technical vocabulary and quiet speech, and they will confidently produce a plausible wrong word rather than an inaudible marker. Our overview of transcribing audio to text online covers the general upload workflow if this is new to you.
One caution: if your ethics approval restricts who may process the audio, check the provider's data handling and deletion policy before uploading anything, and record that check in your data management plan.
Step 6: Set timestamp granularity to match how you will code
Timestamps are how you get back to the audio when a coded extract looks ambiguous, so put them where your unit of analysis is.
For thematic and framework analysis, a timestamp at every speaker turn is usually enough. For interactional analysis, you want them far more frequently, often per utterance. For long unstructured interviews, some researchers use a fixed interval of roughly one every minute or two as a compromise. Whatever you choose, avoid timestamping every line if you plan to export to Word, because dense inline timestamps make the transcript hard to read and can interfere with auto coding by paragraph style. If you want timestamps that link back to media rather than sitting in the text, the same principles apply as when transcribing video with timestamps.
Step 7: Anonymise participants as a separate, deliberate pass
Do not anonymise while you transcribe. You will miss things because your attention is on the words. Finish the transcript, then run a dedicated pass whose only job is de-identification.
Replace direct identifiers first: names, employers, job titles that are unique, schools, street names, specific dates of events. Use consistent placeholders so the text still reads: [large northern city] rather than a blank. Then hunt indirect identifiers, which are what actually re-identify people: an unusual role combined with a sector, a rare medical condition, the number of children plus a location, an anecdote a colleague would instantly recognise.
Do not rely on find and replace alone. It misses nicknames, misspellings and possessives. Read through once with fresh eyes. Keep the identifier key in a separate encrypted location, and decide now whether you will destroy that key at the end of the project or retain it for follow-up.
Step 8: Export the transcript into NVivo, ATLAS.ti or Dovetail
Each package prefers a slightly different route.
NVivo imports Word and plain text files. The payoff for a consistent template comes here: if you format speaker labels using paragraph styles such as Heading 1, NVivo can auto code the whole transcript by speaker in one action, which gives you a case node per participant without manual work. If your interviews follow the same topic guide, styling section headings the same way lets you auto code by question as well.
ATLAS.ti imports Word, plain text and PDF documents, and it also supports linking a transcript to the original media so that clicking a timestamped line jumps to that point in the audio. Timed formats such as SRT or VTT are what enable that sync, which is a good reason to keep those exports even if your written transcript is a Word file.
Dovetail is built around uploading the media itself and attaching or importing a transcript, and it accepts common transcript and subtitle formats. It is the friendlier option for applied UX research where you want highlights and clips alongside codes.
Before you import fifty files, import one and check that speaker labels, timestamps and special characters survived the round trip. Fixing a template once beats fixing fifty transcripts.
Step 9: Set retention, backup and deletion dates now
Write the dates into your data management plan while you remember the detail. Note when raw audio gets deleted, whether anonymised transcripts are retained for longer, where the encrypted identifier key lives, who has access, and which third parties held a copy at any point and when they deleted it. If you promised participants that recordings would be destroyed after transcription, put a calendar reminder in now and actually do it. Retention promises are the part of a consent form most often broken by accident.
Common problems and how to fix them
Overlapping speech turns into nonsense. Automatic systems handle crosstalk poorly. Mark the overlap manually with your chosen notation and transcribe both speakers from the audio; this is one place where the manual method has no substitute.
Accents, dialect and technical vocabulary come back wrong. Build a short list of recurring project terms and participant pseudonyms, then run find and replace for the specific mishearings after each draft. The same misrecognitions repeat across interviews, so the list gets faster to apply.
Room noise ruins the recording. You cannot fix this after the fact, so fix it at source: record on a lapel or handheld microphone rather than a laptop, sit away from air conditioning, and always run a thirty second test before starting. For remote interviews, record locally rather than relying only on the platform stream.
Speaker labels are swapped or merged. Diarisation struggles when two speakers sound similar or one participant dominates. Check the first two minutes of every transcript against the audio, since a label error at the start usually propagates.
The transcript imports into NVivo as one undifferentiated document. This is almost always a styling problem. Apply proper paragraph styles to speaker labels rather than manual bold formatting, and re-import.
Too many inaudible gaps. Try headphones instead of speakers, slow the section to half speed, and boost the volume of that segment. If it stays unclear after two attempts, mark it and move on. Guessing contaminates data.
Frequently Asked Questions
Should I transcribe qualitative interviews myself or use software?
Doing the first pass automatically and correcting it yourself is the best of both. You keep the close familiarity with the data that manual transcription gives you, since correcting requires listening to every second anyway, but you skip the slow mechanical typing. Transcribing entirely by hand is still worth it for conversation analysis, where the notation is too fine grained for any automatic system.
What is the difference between verbatim and intelligent verbatim transcription?
Full verbatim records every sound including fillers, false starts, repetitions and pauses, which matters when how something was said is part of your analysis. Intelligent verbatim keeps every substantive word but removes fillers and stammers for readability, with no paraphrasing. Most thematic and framework analysis uses intelligent verbatim. State which one you used in your methods section either way.
How do I anonymise interview transcripts properly?
Run anonymisation as a separate pass after transcription, replacing direct identifiers with consistent bracketed placeholders and then hunting indirect identifiers such as unusual job roles or recognisable anecdotes. Keep the file mapping pseudonyms to real names encrypted and stored apart from the transcripts. Read the finished transcript through once manually, because find and replace alone always misses nicknames and possessives.
Can I import an SRT or VTT file into NVivo or ATLAS.ti?
ATLAS.ti supports linking timed transcripts to the original media so a coded line can jump straight to that moment in the audio, and Dovetail accepts common subtitle formats alongside the media file. NVivo works primarily from Word and plain text, so convert or paste the transcript into your styled template first. Keeping both a Word version and a timed version of every transcript avoids re-exporting later.
The short version
Decide your verbatim style and notation before you touch the first recording, transcribe the first pass automatically and correct it yourself, anonymise in a separate deliberate pass, and test your import into NVivo, ATLAS.ti or Dovetail with a single file before you process the whole set. The transcription itself is the easy part; the conventions you fix in advance are what make the coding go quickly.