Speech and voice models · try before you read

AI Models

Each entry below runs in the browser first — put your own audio or text through it, then read the write-up if it earned your attention. Models are added as they become available to try.

The collection

What you can try right now

Every model gets its own page: a demo you can drive, what it does well, and what it does not.

Speech Recognition · Reference

gpt-transcribe

OpenAI's July 2026 default for file transcription: $0.0045 per minute, roughly half Whisper's error rate on multilingual audio, plain-text output. The reference page keeps pricing, benchmarks and its output gaps in one place.

OpenAIReleased Jul 2026$0.0045/minReference
Explore gpt-transcribe
Speech Recognition

Nemotron 3.5 ASR

An open 0.6B streaming recogniser from NVIDIA, multilingual and punctuation-aware. Feed it a clip or a URL and watch the text arrive as it listens.

Streaming ASR~40 localesOpen weightsAudio to text
Explore Nemotron 3.5 ASR
Text to Audio

Seed Audio 1.0

ByteDance's Seed Audio 1.0, hosted on fal: audio generated from a written prompt, a reference clip, or an image. Build the API call here and read the responsible-use notes before you ship it.

Text to audioReference audioImage guidancefal API
Explore Seed Audio 1.0
Text to Speech

Miso One

MisoTTS from Miso Labs — open weights, and far more expressive than most TTS. The demo will clone a voice from a short sample so you can judge it yourself.

Emotive TTSVoice cloningOpen weightsEnglish
Explore Miso One
Multilingual Text to Speech

OmniVoice

k2-fsa built this for reach: zero-shot synthesis across 600+ languages. The Hugging Face demo covers cloning and voice design without any training step.

600+ languagesVoice cloningVoice designApache-2.0
Explore OmniVoice
Multilingual Text to Speech

VoxCPM

VoxCPM2 from OpenBMB skips the tokenizer entirely — 30 languages, controllable cloning, voice design, and 48 kHz output. The official demo is linked below.

30 languagesVoice designVoice cloning48 kHz
Explore VoxCPM

The list keeps growing

New models land here once there is something worth trying. In the meantime, the gpt-transcribe reference, Nemotron 3.5 ASR, Seed Audio 1.0, Miso One, OmniVoice and VoxCPM are all live.