About JustTranscribe
JustTranscribe (justtranscribe.ai) turns audio and video into accurate, timestamped text in minutes — WhatsApp voice notes, recorded classes and meetings, podcasts, and public YouTube, TikTok or Instagram videos — with speaker labels, translation into 111 languages and exports to DOCX, PDF, SRT and more. It is built Spanish-first for Latin America and works in English too.
What it is — and what it isn't
One engine behind a set of free tools: upload a file or paste a public link, and the audio is transcribed word by word with timestamps by a Whisper-class speech model, then read by an AI layer that writes a summary and finds the key moments. For social videos, that layer also maps the hook, scenes, pacing and structure, and answers questions grounded in the transcript.
It is automatic transcription. It does not offer human-reviewed transcripts, it can't fetch private or login-only videos (upload the file instead), and it does not transcribe live. Any automatic transcript is a working draft: clear audio comes back close to final; noisy rooms and overlapping voices need a read-through, which the clickable timestamps are there for.
The first transcript is free with no account (up to 40 minutes); after that, a free account raises the limit to 150 minutes and 500 MB per file — unlimited transcripts during the beta, no card needed.
Who is behind it
JustTranscribe is operated from San Sebastián, in the Basque Country (Spain), and is legally represented by Unai Goikoetxea, who also writes and signs the guides on the blog. There is no other "JustTranscribe" behind this site: we are not related to justtranscribe.com or to similarly named mobile apps.
Unai Goikoetxea · Founder and author
Unai runs JustTranscribe from San Sebastián, in the Basque Country, and writes these guides from what people actually upload — WhatsApp audios, recorded classes and meetings, YouTube videos — and from the tool's real limits, not the marketing ones. Spot a mistake? The address in the footer reaches him.
How the transcription works
Speech-to-text runs on OpenAI's Whisper-class model; the summary, the video breakdown and translation run on Anthropic's Claude; speaker detection, only when you turn it on, runs on Deepgram. Files and transcripts are stored on our own infrastructure (AWS) and stay in your library until you delete them — any transcript, or the whole account, at any time.
Long recordings are supported up to 150 minutes and 500 MB per file with a free account, 40 minutes for the first transcript without one. The transcriber listens at 16 kHz mono, so the format and bitrate your recording comes in matter far less than the recording itself.
How we write the guides and comparisons
The blog and the tool pages are written by a person and kept honest by a few rules: no statistics we didn't measure, no prices we didn't read on the competitor's own page on a stated date, and — where we show what the tool produces — a real, unedited output of a public-domain or Creative-Commons recording, with its input file linked so you can run it yourself. Where the model gets something wrong in those samples, the caption says so.
Comparison pages are dated, say what the other tool does better, and carry an address to write to when something is out of date.
Privacy, briefly
Your recordings are used to produce your transcript — nothing else. Analytics are cookieless, and the full list of processors (OpenAI, Anthropic, Deepgram, AWS, Cloudflare, PostHog, the video-fetch providers) with what each receives is in the privacy policy.
Contact
Questions, corrections, a transcript that came back wrong, a price on a comparison page that changed — write to bramontiventures@gmail.com.