GPT Transcribe Online: Speech to Text with Subtitles and Speaker Labels

Searching for gpt-transcribe? That's OpenAI's new transcription model — accurate, cheap, and plain-text only. GPT Transcribe is the browser tool that adds the missing layer: SRT/VTT subtitles, speaker labels, timestamps. Upload audio or video, record live, or paste a link — five free minutes, no API key.

🎬 SRT & VTT subtitles🗣️ Speaker separation🌍 100 languages
Upload ConsoleReady

AI speech to text

Sign in

Audio intake

OpenAI WhisperWhisper DiarizationWord-Level TimestampsSRT · VTT · DOCX · JSON · PDF · TXTAudio + Video Input
The model behind the name · July 2026

What gpt-transcribe Is — and What It Leaves Out

OpenAI shipped gpt-transcribe on July 29, 2026 as its new recommended speech to text model: $0.0045 per minute, roughly 34× faster than real time on files, and on Common Voice across 22 languages it cuts word error rate from 40.37% with whisper-1 down to 19.27%. Independent testing places it at 3.31% word error rate on Artificial Analysis' benchmark — behind ElevenLabs Scribe v2 at 2.3%, but clearly ahead of its predecessor gpt-4o-transcribe at 4.01%. It takes keyword hints and several language hints at once, which helps with product names and code-switching. What comes back, though, is plain text — the subtitle files, speaker labels and word-level timestamps that real work depends on are not part of it.

GPT Transcribe workspace showing an uploaded recording and its transcript

Where gpt-transcribe Is Strong

Accuracy on difficult audio, and context you can steer: keyword hints for product names and acronyms, multiple language hints for conversations that switch language mid-sentence. Hints are suggestions rather than instructions — a keyword only reaches the transcript if the audio actually contains it.

Where It Stops

No SRT or VTT output, no speaker diarization, no word-level timestamps, no translation endpoint. Those still come from whisper-1 or gpt-4o-transcribe-diarize, which is why a caption or interview workflow cannot run on gpt-transcribe alone.

What GPT Transcribe Adds

This site closes that gap in a browser tab. Transcription runs on OpenAI Whisper with diarization, so a one-hour panel comes back split by speaker, timestamped, and ready to export as SRT or VTT — no API key, no integration work, five free minutes to start.

Tool versus endpoint

Why Use This Instead of Calling the API

Calling a transcription endpoint yourself means provisioning a key, watching per-minute billing, handling uploads, polling for results, then writing the code that turns a JSON response into a subtitle file. GPT Transcribe is the finished version of that work.

GPT Transcribe workflow from audio or video upload through to exported transcript

Nothing to Wire Up

No key to provision, no usage dashboard to watch, no upload pipeline to build. Sign in and the first recording is transcribing within seconds — new accounts start with five free minutes.

Transcription use cases across meetings interviews podcasts lectures and captions

Subtitles and Speakers, Out of the Box

SRT and VTT come out ready to drop on a video timeline, and speaker labels arrive already applied. Both are things gpt-transcribe does not produce at all, and both are exactly what interview and captioning work runs on.

Transcript panel in GPT Transcribe with search, editing and export options

Six Export Formats, Not One Response

TXT, SRT, VTT, DOCX, JSON and PDF. A raw API response hands you one shape and leaves every conversion after that to you.

Speech to Text AI Features

What it takes to transcribe audio to text and actually get the result somewhere useful.

Audio and Video in One Slot

MP3, WAV, M4A, FLAC, OGG, MP4, MOV and WebM all go to the same place — audio to text and video to text down one path, with no converting a screen recording first.

Record Without Leaving the Page

Hit record for a stand-up, a call on speaker, or a thought you do not want to lose, and it moves into transcription the moment you stop.

Paste a Media Link

Give GPT Transcribe a URL and it fetches the audio itself — useful when the file is large or lives somewhere you would rather not download from twice.

100 Languages, Auto-Detected

Name the language or let it work one out. Detection matters most on short clips and on recordings that move between languages partway through.

Search and Correct In Place

Jump to any word, fix a misheard name once, and keep the timestamps intact so exported captions stay in sync with the video.

Six Export Formats

TXT, SRT, VTT, DOCX, JSON and PDF — captions, documents, archives and downstream automation all covered without a second tool.

Model comparison

gpt-transcribe vs Whisper: What Each One Actually Does

OpenAI released gpt-transcribe on July 29, 2026 and recommends it for transcription now — on plain text it is both cheaper and considerably more accurate, and it processes files at roughly 34× real time. Its own migration guide, though, lists four reasons to stay on whisper-1 — and three of them are things transcription work usually depends on.

Capabilitygpt-transcribewhisper-1
StatusCurrent default, released July 29, 2026Legacy — kept for timestamps, subtitles and translation
Price per minute$0.0045$0.006 — about 33% more
Word error rate, Common Voice across 22 languages19.27%40.37%
Word and segment timestampsNot availableAvailable — the only OpenAI-hosted option, via timestamp_granularities[]
Native SRT and VTT outputNot availableAvailable through response_format
Translation into EnglishNot availableAvailable on the /v1/audio/translations endpoint
Speaker separationNeeds the separate gpt-4o-transcribe-diarize modelNot native either — needs a diarization stage alongside it
Keyword and multi-language hintsSupported, and it reports the language it detectedSingle legacy language hint only
Streaming transcript eventsSupported, including Realtime committed turnsNot supported

Pricing and capability rows follow OpenAI's model pages and its migration guide, checked July 30, 2026; error-rate figures come from OpenAI's published Common Voice comparison and Artificial Analysis' independent benchmark rather than testing we ran. GPT Transcribe runs OpenAI Whisper with a separate diarization stage — the multi-stage pipeline OpenAI describes for speaker-labelled output — which is why subtitles, timestamps and speaker labels come straight out of this site with nothing to assemble.

GPT Transcribe Pricing

Every account opens with 5 free minutes — enough to run a real recording through GPT Transcribe before paying anything. A subscription covers 1GB uploads, speaker labels, the transcript editor, and every export format, with AI summaries and translation from Pro upward. Credit packs exist for the months you go over, so an occasional long recording doesn't push you onto a bigger plan.

Starter

$9.90$4.90/mo

Steady, low-volume transcribing at the lowest yearly rate.

Includes

  • 1,440 minutes of transcription per year
  • Works out to 120 minutes a month
  • Charged once a year at $58.80
  • Files up to 1GB, uploaded or recorded in-page
  • Timestamps plus per-speaker labelling
  • In-browser transcript editor
  • Six export formats: TXT, SRT, VTT, JSON, PDF, DOCX

Billed for the full year; the figure shown is what it averages per month.

Pro

Best value
$29.90$14.90/mo

The tier most creators and small teams settle on, AI tools included.

Includes

  • 7,200 minutes of transcription per year
  • Works out to 600 minutes a month
  • Charged once a year at $178.80
  • Carries over everything in Starter
  • AI Summary and AI Analytics on any transcript
  • Ask questions about a transcript in chat
  • Translation into 100+ languages
  • Send transcripts, summaries and translations by email

Billed for the full year; the figure shown is what it averages per month.

Max

$49.90$24.90/mo

Heavy transcription loads and AI workflows, at the best per-minute rate.

Includes

  • 36,000 minutes of transcription per year
  • Works out to 3,000 minutes a month
  • Charged once a year at $298.80
  • Carries over everything in Pro
  • Lowest cost per minute at volume
  • Built for putting batches through at once
  • AI Summary, AI Analytics and transcript chat all included
  • Translation into 100+ languages, with email delivery
  • Your audio stored privately

Billed for the full year; the figure shown is what it averages per month.

FAQ

gpt-transcribe: Common Questions

Pricing, accuracy, subtitles, the realtime model, and what this site runs on.

OpenAI's recommended speech to text model, released July 29, 2026, used both for transcribing completed files and for Realtime input. It accepts audio plus optional text context, returns text, supports streaming, and takes keyword hints and multiple language hints. It succeeded gpt-4o-transcribe, the March 2025 model, which OpenAI now marks as not recommended for new integrations.

Click questions to expand detailed answers

Put a Recording Through GPT Transcribe

Upload audio or video, record live, or paste a link. You get a timestamped transcript with speakers separated, exportable as SRT, VTT, DOCX, JSON, PDF or TXT — starting with five free minutes.

Real-timePrivacy FirstNo Setup