Voice platform
Rime
Conversational text-to-speech models for voice agents, available as an API, in your VPC or on-premises.
- Pricing at a glance
- Starter from $0.03 per 1K characters; Enterprise custom
Official Rime signup; the Starter plan includes free usage without a credit card.
Choose a conversational voice for the agent you build
What you get
Stream natural agent speech with pronunciation control into LiveKit, Pipecat or Vapi, or run it in your VPC or on-premises.
What to account for
Rime is the voice only. Reasoning, telephony and actions come from your agent platform, and language coverage is narrower than some rivals.
What the platform provides
- Speech models
- Coda for conversational speech and the Mist family with pronunciation controls, streamed over HTTP or WebSocket.
- Where it runs
- Rime's cloud API, your VPC, or your own NVIDIA hardware on-premises.
Connections & compatibility
- Connects with
- LiveKit, Pipecat, Vapi, Daily
Pricing & terms
Starter is usage-based at $0.03 (Mist v3) to $0.05 (Coda) per 1,000 characters with 20 concurrent streams and free starting usage; the pricing page and docs state different free amounts. Enterprise pricing, voice clones, SLAs and on-prem licensing are custom. Add LLM, speech-to-text and telephony costs from the rest of your stack.
Before committing
Read your own names, numbers and addresses aloud through Coda and Mist v3, and time first audio from your region under load.
- Which of our languages and voices are available on the cloud API versus on-prem?
- What time to first audio should we expect from our region at our peak concurrency?
- How are request text and audio retained, and is any of it used for training?
- What does an enterprise voice clone cost, and what consent do you require?
Controls to confirm
Rime offers SOC 2 reports and a HIPAA BAA, and VPC or on-premises deployment keeps audio and text inside your network. Confirm retention for cloud requests, consent terms for voice clones, and whether your languages are available on-prem, since the packaged language set can differ.
What to consider next
A different approach
Consider Cartesia if you want low-latency speech models plus a managed agent runtime from one vendor.
A different approach
Investigate ElevenLabs for a larger voice library and a hosted agent platform.