Voice platform
Gradium
Low-latency text-to-speech, speech-to-text and speech translation APIs from the Kyutai research team.
- Pricing at a glance
- Free; $13-$1,615/mo credit plans; Enterprise custom
Official plans page; the Free plan includes 45,000 monthly credits for non-commercial use.
Add European-language speech to your own agent stack
What you get
Stream text-to-speech, speech-to-text and speech translation in five European languages, with EU or US data residency.
What to account for
Gradium supplies speech components only, covers five languages, and the free plan excludes commercial use.
What the platform provides
- Speech APIs
- Streaming text-to-speech, speech-to-text with semantic turn detection, and speech-to-speech translation.
- Coverage
- English, French, German, Spanish and Portuguese, with EU or US data residency and voice cloning.
Connections & compatibility
- Connects with
- LiveKit, Pipecat, Vapi, Python SDK
Pricing & terms
Plans use credits: text-to-speech costs 1 credit per character (about 45,000 per hour of speech) and speech-to-text 3 credits per second. Free gives 45,000 credits without commercial use; paid plans run from $13 a month (225,000 credits) to $1,615 a month (45M credits), with extra credits at $3.80-$6.90 per 100,000. Add LLM and telephony costs from your agent stack.
Before committing
Send 8 kHz telephony audio in your target languages and interrupt mid-reply. Check turn detection and time to first audio in your stack.
- How does latency and accuracy hold up on 8 kHz telephony audio in our languages?
- Which plan covers our monthly characters and transcription seconds, and what happens at the limit?
- Is audio or text retained or used for training, and can we keep everything in the EU?
- What consent and verification do you require for cloned voices?
Controls to confirm
Gradium cites ISO 27001 certification and a trust center, and lets you pin sessions to EU or US endpoints. Confirm retention of audio and text, whether submitted feedback or data is used to improve models, and consent controls for voice cloning. Morak has not audited these claims.
What to consider next
A different approach
Consider Deepgram for broader speech-to-text language coverage and a managed voice agent API.
A different approach
Investigate Cartesia for low-latency speech models with a managed agent runtime.