Short answer: upload the recording to a transcription tool that supports speaker labels, turn labels on before you process the file, then expect to manually fix speaker turns wherever participants talk over each other. Focus groups are harder than a two-person interview because you often have four to eight voices, side conversations, and a moderator trying to keep order. AI gets you 70-80% of the way there. The rest is a proofreading pass focused on crosstalk moments and correctly attributing quotes.
Here is the full process, using File Transcribe for the upload and editing steps.
Step 1: Record with speaker separation in mind, if you still can
If you have not recorded yet, a few setup choices save real editing time later:
- Use a central conference mic or multiple mics rather than one laptop mic across the room
- Seat participants so the mic can distinguish voices (avoid clustering everyone on one side)
- Have the moderator ask people to state their name before their first comment, which helps you rename speaker labels faster afterward
If the recording already happened and setup is out of your hands, skip to Step 2. Bad seating and single-mic setups are common in real focus groups; the workflow below still works, it just needs more manual cleanup.
Step 2: Upload the file and turn on speaker labels
Go to File Transcribe and upload the recording, guest upload works for a first pass with no signup. This is the dedicated page for the use case: focus group sessions.
Before processing starts, enable speaker labels. This is not optional for focus groups the way it might be for a single-narrator lecture. Without labels, you get a wall of text with no way to tell who said what, which defeats the purpose of most focus group transcripts.
Step 3: Let the draft process
Processing time depends on the length of the discussion, not the number of speakers directly, though heavily overlapping audio can take slightly longer to process than clean turn-taking. A 60-90 minute focus group is a normal length; budget accordingly for the edit pass that follows.
Step 4: Fix speaker labels first, before anything else
This is the step that separates a usable focus group transcript from a frustrating one. Work through the draft in this order:
- Rename generic labels. Speaker 1, Speaker 2, and so on become actual names or roles (Moderator, Participant A, Participant B) based on context clues in the first few minutes.
- Fix merged turns. When two people talk over each other, the AI sometimes attributes both voices to one speaker, or splits one continuous thought into two speakers. Listen to these sections directly rather than trusting the text.
- Check quiet participants. Someone who speaks less or sits farther from the mic is more likely to get missed segments or wrong attribution. Scan for gaps where you know someone spoke.
- Confirm the moderator's questions are attributed correctly. Researchers often quote moderator prompts alongside participant answers; getting this backwards changes the meaning of a quote.
Step 5: Handle crosstalk sections deliberately
Crosstalk is the single biggest accuracy problem in focus group transcripts, more than any other factor including audio quality. When multiple people respond to a moderator's question at once, or react to another participant's comment, expect the draft to:
- Merge overlapping speech into one speaker's line
- Drop short interjections entirely ("Right," "Exactly," a quick laugh)
- Occasionally flip which speaker said which half of an overlapping exchange
For sections where crosstalk matters to your analysis (agreement, disagreement, energy in the room), listen and manually annotate rather than relying on the transcript alone. For sections where it does not matter (background chatter while someone else has the floor), a rough approximation is usually fine. Decide which sections deserve the extra time before you start editing everything at the same level of care.
Step 6: Export in the format your analysis needs
- Coding in qualitative analysis software (NVivo, Dedoose, etc.): Export TXT or DOCX with speaker labels intact.
- Report quotes for a client deliverable: PDF locks formatting for a clean handoff.
- Video-recorded focus groups needing captions: Export SRT or VTT once you sign in for a free account.
When to bring in more structure
If your focus groups run weekly or are part of a larger research program, a few habits pay off:
- Standardize speaker naming conventions across sessions (Participant A, B, C rather than first names, if anonymization matters)
- Keep a consistent recording setup so accuracy stays predictable session to session
- Build in edit time as part of the research timeline, not an afterthought after transcription "should just work"
For the underlying concept of how speaker splitting works and where it breaks, see what is speaker diarization.
FAQ
How accurate is AI transcription for focus groups?
Lower than a clean two-person interview, because more speakers and more crosstalk both hurt accuracy. Expect a solid first draft on clear turn-taking sections and real editing work on overlapping speech. Budget proofreading time accordingly.
How many speakers can File Transcribe label in one recording?
Speaker labels work across multiple voices in one file. Very large groups (eight or more active speakers) need more manual cleanup than a typical four-to-six person focus group.
What if two participants have similar voices?
Similar voices, especially same gender and similar mic distance, are harder for any speaker labeling system to separate consistently. Expect more manual renaming in these sections, and verify identity by listening rather than trusting the label alone.
Should I record video or audio only for a focus group?
Audio is enough for transcription. Record video only if you need visual context (body language, product demos) for your analysis; it does not improve transcript accuracy on its own.
Can I anonymize speaker names in the transcript?
Yes. After transcription, rename speaker labels to Participant A, B, C instead of real names during the editing step, before you share the transcript outside your research team.
Start your focus group transcript
Upload the recording on the homepage, turn on speaker labels, and budget real editing time for the crosstalk sections. That combination gets you a usable transcript faster than trying to type it by hand.
Related: Focus group sessions · What is speaker diarization · Interview recordings · Pricing
More guides
- Transcription guides
- Try File Transcribe free
- Transcript format guide
- Best transcription software
- Transcribe YouTube videos
