How to Transcribe Long Audio Recordings (1–2 Hours) Without Losing Your Mind
By Unai Goikoetxea · Published
To transcribe long audio, upload the recording to a transcription tool that supports the full length, wait a few minutes, then work with the result by searching and clicking timestamps instead of reading it top to bottom. With JustTranscribe you can paste a link or upload a file up to 500 MB; registered users can process up to 150 minutes, and the first transcript is free with no account for recordings up to 40 minutes. Exports include DOCX, PDF, TXT, SRT, VTT, CSV and Markdown.
A two-hour interview produces a wall of text. Roughly 15,000 to 20,000 words, depending on how fast people speak. Nobody reads that from beginning to end. The trick with long recordings is not the transcription itself, which is now fast, but how you find the parts that matter afterwards.
This guide covers what to expect at 60 to 150 minutes, when and how to split a file, why word-level timestamps beat linear reading, how speaker labels help in group conversations, and a short checklist for recording audio that transcribes cleanly.
What changes when audio gets long
Short clips are forgiving. A three-minute voice note can be wrong in one place and you will still understand it. In a two-hour recording, small problems multiply: one person drifting away from the microphone, a cough during the key sentence, forty minutes of small talk before the interesting part.
Three things become important at this length. First, the tool must accept the whole file, so you do not have to guess where to cut. Second, you need timestamps that are precise enough to jump back into the audio and check a sentence. Third, you need structure — a summary, key moments, or the ability to ask the recording a question — because scrolling is not a search method.
- 60 minutes: roughly 8,000–10,000 words of text
- 90 minutes: a document you will only ever read in fragments
- 120–150 minutes: plan on searching, not reading
Length and size limits, stated honestly
Every tool has a ceiling. Knowing it before you start saves a wasted upload.
In JustTranscribe, a registered user can process files and videos up to 150 minutes. The first transcript, which needs no account at all, accepts up to 40 minutes. File size is capped at 500 MB, and accepted formats include MP3, M4A, WAV, OGG and OPUS (including WhatsApp voice notes), plus MP4 and MOV video. You can also paste a public link from YouTube, TikTok, Instagram, Facebook, Pinterest or a public Google Drive file.
Two limits matter for long recordings in particular. Private or login-only videos cannot be fetched, so a webinar behind a password has to be exported and uploaded as a file. And there is no live transcription: the recording has to exist before you process it.
When to split a file, and how to do it cleanly
If your recording is longer than the limit — a three-hour conference session, a full day of interviews — split it. Splitting is also useful when only part of the recording matters, because you avoid processing an hour of setup noise.
Cut on a natural pause, never mid-sentence. A gap between speakers, a question, the start of a new topic. If you cut mid-word the first sentence of the second part will read badly, and you will spend more time fixing it than you saved.
Free tools that do this well enough: Audacity on desktop (open the file, select a range, export selection), or ffmpeg from the command line if you are comfortable with it. On a phone, most voice recorder apps can trim. Whatever you use, export to MP3 or M4A rather than uncompressed WAV, so the parts stay well under 500 MB.
Name the parts so you can reassemble them later — interview-part1, interview-part2 — and note the offset. If part two starts at 01:30:00 of the original, remember that a timestamp of 00:05:12 in that transcript is really 01:35:12. Write the offset in the file name if you are dealing with more than two parts.
- Split at a pause, not mid-sentence
- Keep 5–10 seconds of overlap so nothing falls between parts
- Export as MP3 or M4A to stay within the size limit
- Record the time offset of each part
Doing it with JustTranscribe
JustTranscribe is a web app, so there is nothing to install. You open justtranscribe.ai, paste a public link or drop the file in, and wait. Processing takes minutes rather than hours, and the language is detected automatically — Spanish, English and many others — so you do not have to declare it in advance. The first transcript is free and needs no account, which is a reasonable way to test the quality of your own audio before committing a two-hour file.
What comes back is not just a block of text. The transcript has word-level timestamps, and each segment is clickable: press it and the audio jumps to that moment. That is the feature that makes long recordings manageable. You search for a name or a phrase, click the hit, and hear the original sentence in context.
On top of the transcript there is AI analysis: a summary, the main idea, key moments, scenes and script structure, and a hook — a quotable line with its timestamp. You can also ask the recording questions and get answers that point back to specific timestamps. For a 90-minute panel discussion, this is usually where you start; the raw transcript is for verification.
When you need the text elsewhere, ask for speaker labels and export. DOCX is the practical choice for editing and sharing with colleagues, PDF for a fixed record, SRT or VTT for subtitles, CSV if you want to sort or filter segments, TXT or Markdown for notes. There is also a share link. Translation into 111 languages is available if the recording needs to reach a different audience. The service is in free beta.
Speaker labels in long conversations
In a one-person recording, labels do not matter. In a two-hour meeting with five participants, they are the difference between a usable document and a puzzle.
Speaker labels are available on demand. They separate the transcript into turns, so you can see who said what and follow a disagreement across twenty minutes. Automatic separation works best when people take turns rather than talking over each other, and when each voice sounds distinct. Expect to correct a few boundaries by hand in a crowded room.
A practical habit: at the start of the recording, ask each person to say their name once. Later you can map Speaker 1 and Speaker 2 to real names in seconds, without listening back to guess.
Stop reading, start searching
The instinct with a fresh transcript is to read it from the top. With 15,000 words that wastes an afternoon.
A better order: read the summary, scan the key moments, then use search for the terms you already know matter — a product name, a date, a person, a number. Click the timestamps to confirm the audio matches the text. Only then, if you are writing something from the recording, read the sections around your hits in full.
If you are looking for something you cannot name precisely, ask the recording a question instead. The answer comes with timestamps, so you can verify it rather than trust it. Treat that as a way to locate passages, not as a replacement for checking them.
Recording-quality checklist for long sessions
Nothing improves a transcript more than better input, and long recordings punish bad setups for hours at a stretch. A few minutes of preparation pays back.
None of this requires equipment you do not have. A phone placed well beats a good microphone placed badly.
- Put the microphone or phone close to the speakers, not in the middle of a large table
- Turn off fans, air conditioning and notification sounds before you start
- Record indoors where possible; wind and traffic are hard to recover from
- Ask people not to talk over each other; one voice at a time transcribes far better
- Do a 30-second test recording and listen to it before the real session
- Keep the device plugged in — a two-hour recording that stops at minute 70 is the worst outcome
Alternatives worth knowing about
Some platforms already give you something for free, and for certain jobs that is enough.
YouTube generates automatic captions for uploaded videos. They are free and adequate for a rough search, but there is no speaker separation, punctuation is inconsistent, and downloading a clean document is awkward. Microsoft Word has dictation and transcription features, and Google Docs has voice typing, but voice typing works on live speech through your microphone rather than on an existing file, which makes it unsuitable for a recording you already have. Most phones now include an on-device voice memo transcription, which is convenient for short notes and often limited in length or language.
The pattern is consistent: built-in tools handle short, single-speaker, clean audio well. Long multi-speaker recordings are where timestamps, speaker labels, search, structured analysis and document exports start to matter.
Step by step
- 1
Check the length and format first
Look at the duration of your recording. If it is under 150 minutes and you have an account, you can process it in one piece. If it is longer, plan the split now. Confirm the file is MP3, M4A, WAV, OGG, OPUS, MP4 or MOV and under 500 MB.
- 2
Split the file if it exceeds the limit
Open the recording in Audacity or any trimming app, find a natural pause near your cut point, select the range and export the selection. Keep a few seconds of overlap between parts. Name the parts in order and note the time offset of each one.
- 3
Open justtranscribe.ai and add your audio
Go to justtranscribe.ai in your browser. Either paste a public link from YouTube, TikTok, Instagram, Facebook, Pinterest or public Google Drive, or drag the file into the upload area. The first transcript is free and needs no account for recordings up to 40 minutes; for longer files, sign in.
- 4
Wait for processing and open the transcript
Processing takes minutes rather than hours, and the language is detected automatically. When it finishes, the transcript appears with word-level timestamps and clickable segments. Play a few seconds from the middle to confirm the audio and text line up.
- 5
Turn on speaker labels for multi-person recordings
Request speaker labels so the transcript is split into turns. Match Speaker 1, Speaker 2 and so on to real names, using the introductions at the start of the recording if you have them. Correct any turn boundaries that look wrong.
- 6
Read the summary and key moments before the full text
Open the AI analysis: summary, main idea, key moments, scenes and the hook with its timestamp. This tells you where the useful parts of a long recording are. Ask the recording a direct question if you are hunting for something specific; the answer comes with timestamps you can check.
- 7
Search for what you need and verify by clicking
Use search for names, dates, numbers and terms you care about. Click each result to jump to that point in the audio and confirm the wording. Fix any misheard names or technical terms directly in the transcript.
- 8
Export in the format you actually need
Choose DOCX for editing and sharing, PDF for a fixed record, SRT or VTT for subtitles, CSV for sorting segments, TXT or Markdown for notes. Use the share link if colleagues only need to read it. Translate the transcript into another language if your audience requires it.
Frequently asked questions
How long can an audio file be to transcribe it in one go?
In JustTranscribe, registered users can process files and videos up to 150 minutes, with a size limit of 500 MB. The free transcript that requires no account accepts up to 40 minutes. Anything longer needs to be split into parts before uploading.
How long does it take to transcribe a two-hour recording?
Minutes, not hours. Automatic transcription does not run in real time, so a long file does not take as long as its duration. Exact time depends on the file and the current load.
Is it better to split a long recording or upload it whole?
Upload it whole if it fits within the limit. A single transcript keeps timestamps consistent and avoids the bookkeeping of offsets. Split only when the recording exceeds the length limit or when just one part of it is relevant to you.
Can I get a Word document from a long transcript?
Yes. DOCX is one of the export formats, along with PDF, TXT, SRT, VTT, CSV and Markdown. Export after you have corrected names and added speaker labels, so you do not repeat the same edits in the document.
Will the tool separate speakers in a long meeting?
Speaker labels are available on demand and split the transcript into turns. Separation is most reliable when people speak one at a time and their voices are distinct. In busy rooms expect to adjust a few boundaries manually.
Can I transcribe a private webinar or a video that requires a login?
Not by pasting the link. Private and login-only videos cannot be fetched. Download or export the recording first, then upload the file directly, as long as it fits the length and size limits.
Try it now
Related articles
- How to Transcribe a Google Drive Video or Audio FileTranscribe a Google Drive video or audio file: set the link to "anyone with the link" and paste it, or download and upload. Timestamps and exports.
- How to Transcribe an Interview With Speaker Labels (and Quote It Safely)A practical guide to transcribe interview audio: recording tips, timestamps, speaker labels, checking quotes against the audio, and clean exports.
- How to Transcribe a Lecture Recording and Turn It Into Study NotesTurn a recorded class into searchable text, a summary and key moments. Steps to transcribe a lecture recording and export study notes as DOCX.