The first decision in transcribing an interview is not which tool to use. It is how faithful the transcript should be. Verbatim keeps every stumble and filler word. Intelligent verbatim removes them. Clean read tidies the grammar as well. Choosing the wrong one wastes hours or produces a transcript that is useless for its purpose.

Most guides skip this and go straight to software recommendations. But a research transcript that has been cleaned up has destroyed the data, and a published quote full of "um" and "you know" makes the speaker look worse than they sounded.

This guide covers the three styles, the conventions for marking what happens in a recording, and a workflow that is faster than typing from scratch.

Choose the style before you start

Three transcription styles and what each is for
Style What it keeps Use it for
Verbatim Everything, including ums, false starts, stutters, repetitions Conversation analysis, legal, linguistics, anything where hesitation is data
Intelligent verbatim Every word spoken, minus filler and false starts Qualitative research, market research, most academic work
Clean read The meaning, with grammar tidied and sentences smoothed Journalism, podcasts, marketing, published quotes

The same thirty seconds of speech in each style:

Verbatim: "So I, um, I think what happened was — well, it's like, the team didn't, they didn't really know, you know, what the plan was."

Intelligent verbatim: "So I think what happened was, the team didn't really know what the plan was."

Clean read: "I think the team didn't really know what the plan was."

Decide before you begin. Converting verbatim down to clean read afterwards is easy. Going the other way is impossible without the audio, and by then you have often deleted the thing you needed.

How to transcribe an interview

  1. Decide on the transcription style before you start, based on what the transcript is for.
  2. Run the recording through an automatic transcription tool to get a first draft.
  3. Play the audio back at reduced speed and correct the draft against it, adding speaker labels.
  4. Mark anything unclear, overlapping or non-verbal using consistent conventions.
  5. Do a final pass checking names, numbers and any quote you intend to publish.

Do not type it from scratch

Typing an interview by hand runs at roughly four to six times the length of the recording. A one-hour interview is most of a working day.

Auto-transcribing first and then correcting runs at roughly one and a half to two times the recording length for clear audio. The machine handles the bulk of the words and you handle the judgement, which is the part it is bad at.

OmniveraLabs' transcription tool takes an audio or video file and returns a punctuated draft, with optional timestamps on every line. Our guide to getting a transcript from a recorded video covers the file-handling side.

The correction pass still matters. An unreviewed automatic transcript is a draft, not a transcript, and the errors it makes are exactly the ones that matter most: names, figures, and technical terms.

Speaker labels

Automatic transcription generally will not tell you who is speaking, so labels are added by hand. Pick one convention and hold it for the whole document.

  • Two speakers: INTERVIEWER: and PARTICIPANT:, or initials if the names are on the record.
  • Research with anonymity requirements: I: for interviewer and P1:, P2: for participants. Keep the key to those codes in a separate file, not in the transcript.
  • Group or focus group: number every participant and label each turn. If a voice is genuinely unidentifiable, UNKNOWN: is more honest than a guess.
  • Formatting: label at the start of the line, followed by a colon, with a blank line between turns. Consistency matters more than which style you pick.

Marking what happens in the recording

These conventions are widely used and worth following, because an editor or supervisor reading your transcript will recognise them.

  • Unclear audio: [inaudible 00:14:32] with the timestamp, so someone can go back and listen. Never guess a word and leave it unmarked.
  • Uncertain word: [unclear: contractor?] when you have a probable reading but are not confident.
  • Overlapping speech: [crosstalk] where two people talk over each other and neither is recoverable.
  • Non-verbal sounds: [laughs], [sighs], [long pause]. Include these only when they carry meaning. In a research interview a pause before an answer is often significant; in a podcast it usually is not.
  • Trailing off: an ellipsis for a sentence the speaker abandons, rather than inventing an ending.
  • Interruptions: an em dash at the cut-off point, followed by the interrupting speaker's turn.
  • Your own notes: square brackets throughout, so anything in brackets is clearly the transcriber's addition rather than something spoken.

Timestamps

How often to timestamp depends on what you will do with the transcript.

  • Every speaker change for interviews you will quote from. Fastest way back to a specific moment.
  • Every few minutes for long recordings you mainly need to navigate.
  • Every line only if the transcript is a step toward subtitles, in which case generate an SRT directly instead.
  • Always next to anything marked inaudible, regardless of the general policy.

Most transcription tools can add timestamps automatically, which is worth turning on even if you later strip them. Adding them by hand afterwards means listening through the whole recording again.

The correction pass

Work through the draft with the audio playing at reduced speed, around 0.75x. Do not read the transcript in silence and fix what looks wrong, because the errors that matter are the ones that read perfectly well.

Priorities, roughly in order:

  • Names. Every person, company and place. A misheard name repeats through the whole document and is the error most likely to embarrass you.
  • Numbers. Figures, dates, percentages, quantities. Automatic transcription is unreliable here and the errors are silent.
  • Technical and specialist vocabulary. Anything specific to the field being discussed.
  • Anything you intend to quote. Listen to those passages twice. A quote is the thing that will be checked.
  • Negations. A dropped "not" reverses a sentence and reads entirely naturally. Worth a specific check on any sentence that surprised you.

Consent, anonymity and storage

Transcription is a data-handling step, not just a typing step, and it is where interview material most often ends up somewhere it should not.

  • Consent covers the transcript too. If the participant agreed to be recorded, be clear whether that included the recording being processed by a third-party service.
  • Check where the audio goes. Cloud transcription means the recording leaves your device. For confidential material, a tool that does not retain uploads, or a local option such as Whisper, is the safer choice.
  • Anonymise in the transcript, not afterwards. Replace identifying details as you correct, and keep the key separately. Retrofitting anonymity to a finished document reliably misses something.
  • Institutional rules take precedence. University ethics approval and newsroom policies often specify where recordings may be stored and for how long. Check before choosing a tool, not after.

Quoting from a transcript

Even a clean-read transcript needs care at the point of publication.

Light tidying of filler and false starts is standard practice and generally uncontroversial. Changing word order, merging separate answers into one quote, or removing a qualifier that softened a claim are not, because they change what the person said rather than how it reads.

If a quote needs a word added for clarity, put it in square brackets. If you cut from the middle, use an ellipsis. Both are recognised conventions that tell a reader the quote has been edited.


Need a first draft to work from? OmniveraLabs' free transcription tool turns an audio or video recording into punctuated text, with optional timestamps on every line.

Frequently asked questions

What is the difference between verbatim and intelligent verbatim?

Verbatim records every sound including filler words, false starts, stutters and repetitions. Intelligent verbatim keeps every word that carries meaning but removes the filler and abandoned starts. Verbatim is used where hesitation itself is data, such as conversation analysis or legal work. Intelligent verbatim suits most research and interview use.

How long does it take to transcribe a one-hour interview?

Typing from scratch takes roughly four to six hours. Auto-transcribing first and correcting the draft takes roughly one and a half to two hours for clear audio. Difficult audio, several speakers, or heavy accents push both figures up considerably.

How do I mark inaudible audio in a transcript?

Write [inaudible] with the timestamp, as in [inaudible 00:14:32], so anyone reading can return to that point in the recording. If you have a probable reading but are not confident, [unclear: word?] is better than either guessing silently or leaving a gap.

Should I include ums and ers in an interview transcript?

Only in verbatim transcription, where hesitation is part of what you are analysing. For research, journalism and most other purposes, removing filler makes the transcript more readable without changing what was said.

How do I label speakers in a transcript?

Put the label at the start of the line followed by a colon, with a blank line between turns. Use INTERVIEWER and PARTICIPANT, initials, or coded labels such as I and P1 where anonymity is required. Keep any key to coded labels in a separate file from the transcript.

Is automatic transcription accurate enough for research?

As a first draft, yes. As a finished transcript, no. Automatic transcription is reliable for clear speech but fails on names, numbers, specialist vocabulary and overlapping speakers, and it produces fluent-reading text even when wrong. A correction pass against the audio is not optional.

Can I transcribe an interview without uploading the recording anywhere?

Yes. Whisper runs locally on your own machine and never sends the audio anywhere, which matters for confidential or ethically sensitive material. It is free and open source, though slower without a GPU.

How much can I edit a quote from a transcript?

Removing filler and false starts is standard and generally accepted. Changing word order, combining separate answers, or cutting a qualifier that softened a claim are not, because they alter what was said rather than how it reads. Mark added words with square brackets and omissions with an ellipsis.

← Back to the blog