JustTranscribe

How to Transcribe an Interview With Speaker Labels (and Quote It Safely)

By Unai Goikoetxea · Published · Updated

To transcribe an interview, record it as a single clean audio file, upload that file to a transcription tool, then add speaker labels and check every quote against the audio before you publish it. With JustTranscribe you can paste a link or upload the recording and get a transcript with word-level timestamps in minutes; the first transcript is free and needs no account. From there you export TXT or DOCX for writing, or SRT and VTT if the interview will appear on video.

The part most people skip is verification. A transcript is a draft of what was said, not proof. Timestamps and clickable segments turn checking a quote into a two-second job instead of a ten-minute hunt through the recording.

This guide covers the recording setup, the transcription itself, speaker labels, how to quote accurately, and what to do when a machine transcript is not enough.

Record for the transcript, not just for your ears

Transcription quality is decided before you press record. A machine hears what your microphone captured, so a few small habits save you hours of correction later.

If you interview two or more people, try to keep voices distinct. A shared microphone in the middle of a table makes speakers sound alike and increases the chance that labels get mixed up.

  • Choose a quiet room. Fans, air conditioning and street noise cost more accuracy than a cheap microphone does.
  • Put the recorder close to the speakers, not next to your laptop.
  • Ask everyone to avoid talking over each other, and repeat questions if they overlap.
  • Say the date, the topic and each person's name at the start. It becomes your reference for labels.
  • Record one continuous file if you can. Multiple small clips make timestamps harder to cite.
  • For remote interviews, record locally as well as in the meeting app; local audio is usually cleaner.

What a usable interview transcript contains

A transcript for journalism or research needs more than the words. It needs a way back to the audio and a clear indication of who spoke.

Word-level timestamps mean each word carries a time position, so you can locate a sentence exactly. That matters when a source says "I never said that" and you need the recording, not your memory.

  • Timestamps at segment level for citation, ideally word level for precision.
  • Speaker labels, corrected to real names once you know them.
  • Searchable text, so you can jump to a keyword instead of scrolling.
  • A stable copy of the original audio, stored with the transcript.
  • An export format that matches the next step: DOCX or TXT for writing, SRT or VTT for video, CSV for coding and analysis.

Transcribing the interview with JustTranscribe

Open justtranscribe.ai in a browser. You can upload the recording (MP3, M4A, WAV, OGG or OPUS, including WhatsApp voice notes, and MP4 or MOV up to 500 MB) or paste a public link if the interview is already published on YouTube, TikTok, Instagram, Facebook or a public Google Drive folder. It cannot open private or login-only videos, so for those you upload the file.

The first transcript is free and needs no account, and that free transcript accepts recordings up to 40 minutes. Registered users can process files up to 150 minutes, which covers most long-form interviews. The language is detected automatically, so a Spanish interview does not need any setting change.

Processing takes minutes. You get a transcript with word-level timestamps where every segment is clickable: click a line and the audio jumps to that point. Ask for speaker labels when you need them. The AI analysis panel gives a summary, the main idea, key moments and a suggested hook with its timestamp, and you can ask the recording direct questions, such as what the source said about funding, and get answers with timestamps attached.

If the interview will be published for another audience, the transcript can be translated into 111 languages. Exports include TXT, SRT, VTT, CSV, Markdown, PDF and DOCX, plus a share link for an editor or a co-author. The service is in free beta. It does not do live transcription, so it is for recordings, not for a session in progress.

Honest alternatives and their limits

No single tool is right for every job, and some options cost nothing at all.

Microsoft Word for the web has a transcription feature for Microsoft 365 subscribers: it produces speaker segments you can insert into a document, but it depends on your subscription and has upload limits. Google Docs voice typing is free, yet it only transcribes what your microphone hears live, without timestamps or speaker separation, so it works for dictating notes rather than for interviews. YouTube auto-captions are free once a video is uploaded, and you can copy the transcript, but punctuation is weak, speakers are not separated, and you have to publish or at least upload the file first.

Manual transcription with a foot pedal and a player remains the most accurate method for difficult audio: heavy accents, poor recordings, legal depositions. It is also the slowest: the usual rule of thumb among professional transcriptionists is four to six hours of typing per hour of audio. A sensible compromise is a machine transcript as the first draft, then manual correction of the passages you plan to quote.

Professional human transcription services exist and are worth it when accuracy is contractual, for example in court or clinical research. Whatever you choose, keep the original audio.

Getting speaker labels right

Automatic speaker separation gives you generic labels such as Speaker 1 and Speaker 2. Your job is to map them to names and to fix the boundaries where the system split a turn incorrectly.

Use the opening of the recording, where you introduced everyone, to identify each voice. Then scan the places where speakers interrupt each other; that is where labels most often swap. Group interviews with four or more participants are the hardest case, and sometimes the honest solution is to label an unclear turn as unidentified rather than to guess.

  • Replace generic labels with real names early, before you start pulling quotes.
  • Check every handover point where one speaker interrupts another.
  • Mark inaudible passages clearly instead of inventing words.
  • If two voices are genuinely similar, note the ambiguity in the file for your editor.

Quoting safely: verify before you publish

Treat every quote as unverified until you have heard it in the recording. Click the segment, listen to the sentence and the sentence before it, and confirm both the wording and the context. Word-level timestamps make this fast, which is the whole point.

Decide your quoting convention and apply it consistently. Clean verbatim removes filler words, false starts and stutters without changing meaning; strict verbatim keeps them and is standard in qualitative research and legal work. Never merge two separate answers into one quote, and never move a clause to make a sentence stronger.

Keep the timestamp next to each quote in your working document. When an editor, a fact-checker or the source asks, you go straight to the second in question. If you send the source a quote to confirm, send the exact wording you will print, not a paraphrase.

  • Listen to each quote in the audio, not only in the text.
  • Include the surrounding sentences when you judge context.
  • Use square brackets for added words and ellipses for removals.
  • Store audio, transcript and quote list together so the trail survives.
  • Note consent: say on the recording that the person agreed to be recorded, and follow local rules.

Step by step

  1. 1

    Prepare the recording

    Make one continuous audio file in a quiet room, with the microphone near the speakers. At the start, say the date, the topic and the name of each participant, and confirm out loud that they agree to be recorded.

  2. 2

    Open JustTranscribe and add the interview

    Go to justtranscribe.ai. Drag in the audio or video file, or paste a public link if the interview is already online. The first transcript is free and needs no account, up to 40 minutes; register for files up to 150 minutes. Files can be up to 500 MB.

  3. 3

    Request speaker labels and start processing

    Turn on speaker labels before you process, then start. The language is detected automatically. Wait a few minutes; you do not need to keep the tab in focus the whole time.

  4. 4

    Rename the speakers

    Open the transcript and use the introduction at the beginning of the recording to identify each voice. Replace Speaker 1 and Speaker 2 with the real names, and check the turns where people interrupted each other.

  5. 5

    Find the passages you need

    Search a keyword to jump to it, or use the analysis panel to see key moments and a summary. You can also ask the recording a question, for example what the source said about the budget, and get an answer with a timestamp.

  6. 6

    Verify each quote against the audio

    Click the segment that contains the quote. The audio jumps to that exact point. Listen to the quote plus the sentence before and after, then correct any word that differs from what you hear.

  7. 7

    Export in the format you need

    Export DOCX or TXT for writing and archiving, CSV for coding and analysis, PDF for sharing with a source, and SRT or VTT if the interview will appear as subtitles on video. Keep timestamps in your working copy and remove them in the clean version you hand over.

  8. 8

    Archive the file and the trail

    Store the original audio, the corrected transcript and your quote list with timestamps in the same folder. Use a share link if an editor or co-author needs access to the transcript.

Frequently asked questions

How long does it take to transcribe a one-hour interview?

Manual typing usually takes four to six hours per hour of audio, depending on audio quality and how many speakers there are. Automatic transcription takes minutes, and then you spend your time on correction and verification instead of typing. Most of the remaining work is fixing speaker labels and checking the sentences you plan to quote.

Can I transcribe an interview for free?

Yes, several ways. On JustTranscribe the first transcript is free and needs no account, for recordings up to 40 minutes. Google Docs voice typing is free but only handles live audio without timestamps, and YouTube auto-captions are free once a video is uploaded, without speaker separation.

How do I separate speakers in an interview transcript?

Ask the tool for speaker labels when you process the file; it will return generic labels such as Speaker 1 and Speaker 2. Then map those to real names using the introduction at the start of the recording. Check the points where speakers interrupt each other, because that is where labels usually get swapped.

Should I include timestamps in the final transcript?

Keep them in your working copy, because they are what let you verify a quote quickly. Remove them in the clean version you send to a source or publish as a text. Exporting twice, once with timestamps and once without, is the simplest approach.

Is an automatic transcript accurate enough to quote from?

Treat it as a first draft. Machine transcription is reliable on clear audio, but names, technical terms and overlapping speech need checking. Listen to each quote in the recording before publishing, and correct the wording to match what you hear.

Can I transcribe an interview recorded on WhatsApp or a phone?

Yes. WhatsApp voice notes in OGG or OPUS format and phone recordings in M4A, MP3 or WAV can be uploaded directly, as can MP4 and MOV video up to 500 MB. Quality still matters, so record in a quiet place and keep the phone close to the speakers.

Try it now

Related articles

← All articles