Upload a few seconds of speech and generate natural narration in that voice — or pick from 26 built-in voices. No third-party API, no per-character billing, nothing leaving your server.
Free to use · Voice samples stay in your browser
The whole pipeline runs on one machine you control. No keys, no quotas, no audio sent to anyone else.
Drop in a short recording — or record straight from your microphone — and the model matches the timbre immediately. No training step, no waiting.
A curated built-in library spanning several source datasets, plus French, German, Italian, Spanish and Portuguese timbres.
Kyutai's Pocket TTS is distilled for small machines. A sentence is typically ready in two to three seconds on plain CPU — no GPU anywhere.
Voice samples and generated audio are held in your browser. The server keeps accounts and usage counts — never your audio.
Every clip is repaired to a valid 24 kHz WAV before it reaches you, so durations, seeking and downstream editors all behave.
Waveform playback with scrubbing and speed control, a searchable voice library, and a full history of everything you've generated.
Upload five to thirty seconds of clean speech, record yourself in the browser, or skip straight to a built-in voice.
Type or paste anything up to a few thousand characters. Pick the voice you want from the library.
Audio streams back within seconds. Play it, scrub it, adjust the speed, and download a clean WAV.
Available the moment you sign in, alongside any voice you clone yourself.
Create an account and the full studio is yours — no card, no quota, no third party.
Create your account