The first decision in transcribing an interview is not which tool to use. It is how faithful the transcript should be. Verbatim keeps every stumble and filler word. Intelligent verbatim removes them. Clean read tidies the grammar as well. Choosing the wrong one wastes hours or produces a transcript that is useless for its purpose.
Most guides skip this and go straight to software recommendations. But a research transcript that has been cleaned up has destroyed the data, and a published quote full of "um" and "you know" makes the speaker look worse than they sounded.
This guide covers the three styles, the conventions for marking what happens in a recording, and a workflow that is faster than typing from scratch.
Choose the style before you start
| Style | What it keeps | Use it for |
|---|---|---|
| Verbatim | Everything, including ums, false starts, stutters, repetitions | Conversation analysis, legal, linguistics, anything where hesitation is data |
| Intelligent verbatim | Every word spoken, minus filler and false starts | Qualitative research, market research, most academic work |
| Clean read | The meaning, with grammar tidied and sentences smoothed | Journalism, podcasts, marketing, published quotes |
The same thirty seconds of speech in each style:
Verbatim: "So I, um, I think what happened was — well, it's like, the team didn't, they didn't really know, you know, what the plan was."
Intelligent verbatim: "So I think what happened was, the team didn't really know what the plan was."
Clean read: "I think the team didn't really know what the plan was."
Decide before you begin. Converting verbatim down to clean read afterwards is easy. Going the other way is impossible without the audio, and by then you have often deleted the thing you needed.
How to transcribe an interview
- Decide on the transcription style before you start, based on what the transcript is for.
- Run the recording through an automatic transcription tool to get a first draft.
- Play the audio back at reduced speed and correct the draft against it, adding speaker labels.
- Mark anything unclear, overlapping or non-verbal using consistent conventions.
- Do a final pass checking names, numbers and any quote you intend to publish.
Do not type it from scratch
Typing an interview by hand runs at roughly four to six times the length of the recording. A one-hour interview is most of a working day.
Auto-transcribing first and then correcting runs at roughly one and a half to two times the recording length for clear audio. The machine handles the bulk of the words and you handle the judgement, which is the part it is bad at.
OmniveraLabs' transcription tool takes an audio or video file and returns a punctuated draft, with optional timestamps on every line. Our guide to getting a transcript from a recorded video covers the file-handling side.
The correction pass still matters. An unreviewed automatic transcript is a draft, not a transcript, and the errors it makes are exactly the ones that matter most: names, figures, and technical terms.
Speaker labels
Automatic transcription generally will not tell you who is speaking, so labels are added by hand. Pick one convention and hold it for the whole document.
- Two speakers:
INTERVIEWER:andPARTICIPANT:, or initials if the names are on the record. - Research with anonymity requirements:
I:for interviewer andP1:,P2:for participants. Keep the key to those codes in a separate file, not in the transcript. - Group or focus group: number every participant and label each turn. If a voice is genuinely unidentifiable,
UNKNOWN:is more honest than a guess. - Formatting: label at the start of the line, followed by a colon, with a blank line between turns. Consistency matters more than which style you pick.
Marking what happens in the recording
These conventions are widely used and worth following, because an editor or supervisor reading your transcript will recognise them.
- Unclear audio:
[inaudible 00:14:32]with the timestamp, so someone can go back and listen. Never guess a word and leave it unmarked. - Uncertain word:
[unclear: contractor?]when you have a probable reading but are not confident. - Overlapping speech:
[crosstalk]where two people talk over each other and neither is recoverable. - Non-verbal sounds:
[laughs],[sighs],[long pause]. Include these only when they carry meaning. In a research interview a pause before an answer is often significant; in a podcast it usually is not. - Trailing off: an ellipsis for a sentence the speaker abandons, rather than inventing an ending.
- Interruptions: an em dash at the cut-off point, followed by the interrupting speaker's turn.
- Your own notes: square brackets throughout, so anything in brackets is clearly the transcriber's addition rather than something spoken.
Timestamps
How often to timestamp depends on what you will do with the transcript.
- Every speaker change for interviews you will quote from. Fastest way back to a specific moment.
- Every few minutes for long recordings you mainly need to navigate.
- Every line only if the transcript is a step toward subtitles, in which case generate an SRT directly instead.
- Always next to anything marked inaudible, regardless of the general policy.
Most transcription tools can add timestamps automatically, which is worth turning on even if you later strip them. Adding them by hand afterwards means listening through the whole recording again.
The correction pass
Work through the draft with the audio playing at reduced speed, around 0.75x. Do not read the transcript in silence and fix what looks wrong, because the errors that matter are the ones that read perfectly well.
Priorities, roughly in order:
- Names. Every person, company and place. A misheard name repeats through the whole document and is the error most likely to embarrass you.
- Numbers. Figures, dates, percentages, quantities. Automatic transcription is unreliable here and the errors are silent.
- Technical and specialist vocabulary. Anything specific to the field being discussed.
- Anything you intend to quote. Listen to those passages twice. A quote is the thing that will be checked.
- Negations. A dropped "not" reverses a sentence and reads entirely naturally. Worth a specific check on any sentence that surprised you.
Consent, anonymity and storage
Transcription is a data-handling step, not just a typing step, and it is where interview material most often ends up somewhere it should not.
- Consent covers the transcript too. If the participant agreed to be recorded, be clear whether that included the recording being processed by a third-party service.
- Check where the audio goes. Cloud transcription means the recording leaves your device. For confidential material, a tool that does not retain uploads, or a local option such as Whisper, is the safer choice.
- Anonymise in the transcript, not afterwards. Replace identifying details as you correct, and keep the key separately. Retrofitting anonymity to a finished document reliably misses something.
- Institutional rules take precedence. University ethics approval and newsroom policies often specify where recordings may be stored and for how long. Check before choosing a tool, not after.
Quoting from a transcript
Even a clean-read transcript needs care at the point of publication.
Light tidying of filler and false starts is standard practice and generally uncontroversial. Changing word order, merging separate answers into one quote, or removing a qualifier that softened a claim are not, because they change what the person said rather than how it reads.
If a quote needs a word added for clarity, put it in square brackets. If you cut from the middle, use an ellipsis. Both are recognised conventions that tell a reader the quote has been edited.
Need a first draft to work from? OmniveraLabs' free transcription tool turns an audio or video recording into punctuated text, with optional timestamps on every line.