How to Transcribe Focus Group Recordings Accurately
A focus group is the hardest audio you will ever hand to a transcription tool. Eight people, one table, everyone talking over the moderator, and half the room too far from the mic. This guide walks through how to transcribe focus group recordings end to end, from mic placement on the day through to a labelled transcript, per speaker talk time, and a themed findings summary you can actually put in front of a client.
Step 1: Set Up the Room So the Transcript Has a Chance
Transcription quality is decided before anyone speaks. No model can separate two voices that arrived at the mic as one blurred signal, so spend ten minutes on the room.
- Use more than one mic. A single omnidirectional boundary mic in the centre of a table sounds fine to a human ear but flattens everyone into the same acoustic space. Two or three boundary mics spaced along the table, or a small lavalier per participant, give a diarisation model the volume and timing differences it needs to tell voices apart.
- Record each mic to its own track. This is the single highest value decision you can make. If your recorder or interface supports multitrack (most field recorders and any DAW will), one track per participant means you do not need diarisation at all; each track is a speaker.
- Seat alternating voices apart. Two similar sounding voices sitting next to each other is the most common cause of swapped labels. Split them across the table.
- Kill the room noise. Air conditioning, a projector fan, a laptop on the table. Turn off what you can, and move the mic away from what you cannot.
- Record a roll call. Before the discussion starts, have every participant say their name and one full sentence, in seating order. This thirty second clip is what you will use later to attach real names to generic speaker labels, and it is worth more than any post processing trick.
Step 2: Record a Clean Master File
Export a single WAV or high bitrate MP3 of the full session, plus the individual tracks if you recorded multitrack. Keep the file uncompressed if you can; heavy compression smears the fricatives and sibilance that models use to distinguish speakers.
If the group ran on Zoom, Teams, or Google Meet, use the platform's own separate audio file per participant option before you record, not after. It is buried in the recording settings and it turns a diarisation problem into a non problem. The same principle applies to any remote session, and the setup detail is covered in more depth in the guide to transcribing Google Meet recordings.
Step 3: Transcribe the Recording Manually (the Free Baseline Method)
You do not need paid software to get a usable focus group transcript. The manual method is slow but it is genuinely the most accurate option for short sessions.
- Open the audio in a free editor such as Audacity, or a media player with variable speed like VLC.
- Set playback to around 0.7x speed. Slower than that and the pitch artefacts make speech harder, not easier, to parse.
- Open a document beside it and type using a fixed speaker convention from the first line:
M:for moderator,P1:throughP8:for participants. Consistency matters more than which convention you pick. - Insert a timestamp every two minutes, or at every topic change.
[00:14:30]style markers let you jump back to verify a quote later. - Use keyboard shortcuts to pause and rewind so your hands never leave the keyboard. In VLC, Shift plus Left arrow steps back a few seconds.
- Mark anything you cannot make out as
[inaudible 00:22:15]rather than guessing. A guessed quote in a client report is worse than a gap.
Budget roughly four to six hours of typing per hour of multi speaker audio. That is why almost everyone moves to an automated first pass and edits from there.
Step 4: Run an Automated First Pass With Speaker Diarisation
Automated transcription gets you to a full draft in minutes, then you spend your time correcting rather than typing. What you want from the tool is a transcript with speaker turns already split, timestamps on every turn, and an export you can edit.
This is where Tapescribe fits the workflow: upload the session file or paste a link, and you get a timestamped transcript with speaker separation plus TXT, SRT, and VTT exports, so you can move straight to labelling instead of typing from scratch. If you recorded multitrack, run each participant track through separately and merge by timestamp afterwards; that gives you perfect speaker attribution with no diarisation guesswork at all.
Step 5: Understand Why Diarisation Degrades With Six or More Voices
It helps to know what the model is actually doing, because it tells you where to look for errors.
Diarisation works by turning short slices of audio into voice embeddings, then clustering those embeddings into groups it believes are individual speakers. With three or four distinct voices, the clusters sit far apart and the job is easy. As you add participants, several things go wrong at once:
- The clusters crowd together. More voices in the same acoustic space means smaller gaps between clusters, and any two similar voices, often same gender and similar age, start bleeding into each other.
- Speaker turns get shorter. Big groups produce rapid back and forth and one word interjections. A half second "yeah" is often too short to produce a reliable embedding, so it gets attached to whoever spoke around it.
- Crosstalk has no correct answer. When two people overlap, the audio contains both voices in one slice. Most systems assign that slice to a single speaker, so the quieter voice simply disappears from the transcript.
- Speaker count estimation drifts. If the tool is guessing how many speakers exist, one person who changes volume between a quiet aside and an emphatic point can be split into two labels, while two quiet participants get merged into one.
Practically: expect a clean result up to about five voices, and expect to do real correction work above that. If your tool lets you specify the expected number of speakers, always set it. Guessing is where most of the error comes from. The same constraints apply to any busy multi speaker recording, which is why webinar recordings with a panel are easier than a focus group despite being longer.
Step 6: Label Participants After the Fact
You now have a transcript reading "Speaker 1" through "Speaker 7". Turning that into real names is a mechanical process.
- Start from the roll call. Play the first thirty seconds and note which generic label was assigned to each name as they introduced themselves. Write the map down: Speaker 3 = Priya, Speaker 4 = Tom.
- Verify each label at three separate points. Jump to a turn near the start, middle, and end of the session for each speaker. Diarisation errors are rarely uniform; a label often holds for twenty minutes then swaps after a break or a chair change.
- Find and replace, longest first. Replace
Speaker 10beforeSpeaker 1, otherwise a naive replace turns "Speaker 10" into "Priya0". - Use pseudonyms if you promised confidentiality. Replace with
P1 (female, 34, urban)style labels rather than names, and keep the real name mapping in a separate access controlled file. - Fix crosstalk gaps as you go. Where you see a turn that reads like two people merged, split it into two lines and listen back to attribute each half. If the second voice is genuinely unrecoverable, mark it
[overlapping speech].
Step 7: Export Per Speaker Talk Time
Talk time is the fastest way to spot a dominated session, and clients ask for it constantly. If a transcript has timestamps on every speaker turn, calculating it is arithmetic.
Export the transcript as SRT or VTT, since both formats give every segment an explicit start and end time. Then for each speaker, sum the duration of their segments and divide by total session length. A short script or a spreadsheet handles this: paste the segments in, subtract start from end per row, and pivot by speaker label.
What to read from the result: any participant below roughly five percent of participant talk time did not really contribute, and you should be cautious about generalising from a quote they never expanded on. If the moderator is taking a large share, the guide was running the room rather than the room running itself. Both are findings worth noting in the write up, not just quality checks.
Step 8: Turn the Transcript Into a Themed Findings Summary
A raw transcript is not a deliverable. The output your stakeholders want is themes with evidence.
- Read once without highlighting. Get the shape of the session before you start coding it.
- Code on the second pass. Tag every meaningful passage with a short label describing what it is about:
price sensitivity,onboarding confusion,trust in reviews. Keep labels short and reuse them. - Cluster codes into themes. Group related codes and name the group in the participants' own language where possible.
- Count participants, not mentions. A theme raised by five different people is stronger than one person mentioning it five times. Note both numbers for each theme.
- Pull two or three verbatim quotes per theme, each with a speaker label and timestamp so anyone can verify it against the recording.
- Write the summary as theme, evidence, implication. One paragraph each. Flag disagreement inside a theme explicitly; the dissent is usually more useful than the consensus.
If you also have chapter markers or topic breaks from the transcription pass, use them as a starting skeleton for your themes; they map closely to how the moderator's discussion guide actually ran.
Common Problems and How to Fix Them
Two participants keep getting merged into one label. Almost always similar voices seated close together. Fix manually using the roll call map, and next time seat them apart or mic them separately.
One participant is split across three labels. Usually caused by volume swings or the tool overestimating speaker count. Set the expected speaker number explicitly on a rerun, or merge the labels with find and replace.
Whole passages of crosstalk are missing. Overlapping speech assigned to a single speaker. Listen back at reduced speed and transcribe those passages by hand; there is no automated fix.
The far end of the table is unintelligible. A mic placement problem you cannot solve in post. Some noise reduction and a volume boost may recover a little, but plan to add a second mic next session.
Names, brands, and jargon are consistently wrong. Build a find and replace list of the terms your study uses and run it across the transcript before you start coding. It takes two minutes and removes a whole class of error.
The file is too long to process in one go. Split the audio at a natural break, transcribe each part, then merge and re-offset the timestamps. Long session handling is the same problem covered in the notes on transcribing lecture recordings.
Frequently Asked Questions
How long does it take to transcribe a focus group recording?
Typing a multi speaker session manually takes roughly four to six hours per hour of audio, because you constantly rewind to resolve overlaps. An automated first pass produces a draft in minutes, and correction work then typically takes between one and two hours per hour of audio depending on how many participants there were. Larger groups sit at the higher end of that range.
Can AI transcription handle eight people talking at once?
It handles eight people talking in turn reasonably well, and eight people talking at once poorly. Genuine overlapping speech usually gets assigned to whichever voice is loudest in that slice, so the quieter contribution is lost. The only reliable fix is recording each participant to a separate track, which removes the separation problem entirely.
Do I need a separate mic for every focus group participant?
You do not need one, but it is the biggest single upgrade to transcript quality. If individual lavaliers are impractical, two or three boundary mics spaced along the table, recorded to separate tracks, get you most of the benefit. A single centre mic will work for a small group of four or fewer.
How do I anonymise a focus group transcript?
Replace speaker names with coded labels such as P1 and P2 during the labelling step, and keep the mapping of codes to real identities in a separate file with restricted access. Also scan the body text for identifying details participants mentioned about themselves, such as employers, locations, or health information, and redact those too. Doing this at the labelling stage is far quicker than retrofitting it after coding.
The work that determines whether a focus group transcript is any good happens on recording day: separate tracks, sensible mic placement, and a thirty second roll call. Get those right and the rest is a fast automated pass followed by an hour of targeted correction. Get them wrong and no amount of processing will recover what the mics never captured cleanly.