Grok Voice Think Fast 2.0
Grok Voice Think Fast 2.0 is xAI’s flagship voice model for building real-time assistants, phone agents, and interactive voice systems that stream audio and text bidirectionally over WebSocket. Developers can configure system instructions, high or no reasoning effort, built-in or custom voices, automatic server-side voice activity detection, silence duration, idle re-engagement, playback speed, and session resumption after temporary disconnects. It accepts PCM, G.711 μ-law, G.711 A-law, and Opus audio through JSON or raw binary frames, with configurable PCM sample rates from telephone quality to 48 kHz. It supports more than 20 languages with native-quality accents, automatic language detection, natural responses in the speaker’s language, and seamless code-switching. Language hints and up to 100 key terms improve transcription of regional speech, names, products, codes, addresses, and specialized terminology, while pronunciation replacements correct spoken output.
Learn more
Cartesia Sonic-3.5
Sonic 3.5 is Cartesia’s fastest, most natural text-to-speech model, built for expressive, real-time voice generation with sub-90ms latency and native support for 42 languages. It is designed to follow transcripts faithfully, voice confirmation codes, and heteronyms correctly without preprocessing, and stay expressive enough to carry a real conversation. It supports languages intended to deliver native-quality speech. Sonic 3.5 focuses on clean audio across every language and voice, with no artifacts to edit out, making it practical for production voice experiences where quality, speed, and consistency matter. Its expressive conversational delivery provides strong pacing and real emotional range, tuned for support and agent transcripts. Alphanumerics such as order numbers, phone numbers, IDs, and emails are spoken naturally in every language, while context-aware English pronunciation helps words like read, bass, and bow land correctly from the surrounding text.
Learn more
Grok Text to Speech (TTS)
Grok Text to Speech (TTS) is a standalone audio API built to help developers generate fast, natural, and expressive speech from text. Built on the same stack that powers Grok Voice, Tesla vehicles, and Starlink customer support, the API makes it straightforward to integrate high-quality voice generation into applications such as voice agents, accessibility tools, podcasts, assistants, customer experiences, and interactive audio products. Grok TTS can turn long-form text into speech through a REST API or generate speech in real time through a WebSocket API, giving developers flexibility for both batch audio generation and live conversational experiences. It is designed around expressive delivery, not just flat narration, with fine-grained control through simple inline and wrapping speech tags. Developers can add natural prosody and emotion using tags, allowing lifelike delivery without complex markup.
Learn more
Cartesia Sonic-3.6
Sonic is a real-time text-to-speech model built for voice agents, combining natural delivery, sub-90ms latency, and native support for more than 40 languages. It is designed to make voice interactions feel effortless, with tone that adjusts to context, consistent pacing, and speech that follows the natural rhythm of conversation. By default, Sonic interprets the emotional subtext of a transcript and calibrates delivery automatically, while non-verbal expressions such as laughter can be inserted directly into the text. The model follows transcripts faithfully, produces clean audio across languages and voices, and handles alphanumeric content such as order numbers, phone numbers, IDs, and email addresses naturally without preprocessing. Context-aware pronunciation helps heteronyms sound correct from surrounding words, while custom pronunciation dictionaries let teams define how proper nouns and domain-specific terms should be spoken.
Learn more