Real-time voice AI models and APIs for voice agents: text-to-speech, speech-to-text, live translation, voice cloning and voice design.

About Gradium

Gradium builds the voice models that sit under voice agents and sells them as streaming APIs. A developer gets text-to-speech whose fastest model starts speaking in under 50 ms, speech-to-text that tells from the meaning when a speaker has finished, live speech-to-speech translation that keeps the speaker's voice, instant voice cloning from ten seconds of audio, voices designed from a text prompt, and a text-to-speech model small enough to run offline on a phone or a CPU. Gradbot, its open-source framework, turns these into a working voice agent in about fifty lines of code, and integrations exist for LiveKit and Pipecat. The models cover English, French, German, Spanish and Portuguese. They run on Gradium's cloud with EU or US data residency and zero data retention, on dedicated instances, or self-hosted. The company was founded in Paris in 2025 by the researchers behind the Kyutai lab and has raised 100 million dollars. It is used by teams building voice agents for customer service, translation, media and consumer applications.

Listed inGenerative Media & Voice AIasVoice Generation Tools, Speech-to-Text APIs

How teams use Gradium

No success cases yet.

Media

No media yet
Screenshots and product tours appear here once the vendor claims this page.

Alternatives & competitors