AI Models
Each entry below runs in the browser first — put your own audio or text through it, then read the write-up if it earned your attention. Models are added as they become available to try.
What you can try right now
Every model gets its own page: a demo you can drive, what it does well, and what it does not.
Nemotron 3.5 ASR
An open 0.6B streaming recogniser from NVIDIA, multilingual and punctuation-aware. Feed it a clip or a URL and watch the text arrive as it listens.
Seed Audio 1.0
ByteDance's Seed Audio 1.0, hosted on fal: audio generated from a written prompt, a reference clip, or an image. Build the API call here and read the responsible-use notes before you ship it.
Miso One
MisoTTS from Miso Labs — open weights, and far more expressive than most TTS. The demo will clone a voice from a short sample so you can judge it yourself.
OmniVoice
k2-fsa built this for reach: zero-shot synthesis across 600+ languages. The Hugging Face demo covers cloning and voice design without any training step.
VoxCPM
VoxCPM2 from OpenBMB skips the tokenizer entirely — 30 languages, controllable cloning, voice design, and 48 kHz output. The official demo is linked below.
The list keeps growing
New models land here once there is something worth trying. In the meantime, Nemotron 3.5 ASR, Seed Audio 1.0, Miso One, OmniVoice and VoxCPM are all live.
Here for transcription instead?
The models above sit alongside the transcription workspace this site is built around.
