Streaming text-to-speech

Streaming text-to-speech that plays as it generates, in the browser or over the phone, with a first byte you can measure.

Definition

Streaming text-to-speech returns audio in chunks as it is generated, over a single HTTP connection. Your app decodes and plays the bytes as they arrive, so the listener hears the first sound almost immediately instead of waiting for the whole clip to render. Real-time text-to-speech is what that makes possible: speech fast enough for a live app, where someone is waiting to hear the answer.

Measured latency

SpeechifyAI’s streaming endpoint returns its first byte in 56 ms at the median and 102 ms at the 90th percentile on Simba 3.2, measured on the production US East path on 15 Sep 2026. The site rounds that to <100ms; the number is a measured median, not a guarantee.

Two independent benchmarks time Simba 3.2 from their own runners to the first sound: Coval measured a 106 ms median time to first audio over the 24 hours to 24 Sep 2026, and Voice Arena read 123 ms p50 on its US English board the same day. Their readings include the network to each runner and any silence before the first sound, so they read higher than our own first byte, and neither is its source. Our reviewed snapshots: Coval, Voice Arena.

How streaming works

You call POST /v1/audio/stream. The response is chunked HTTP: audio bytes flow back continuously. Your client reads the stream, decodes each chunk, and feeds it to the audio output. Playback starts on the first chunk and continues as generation continues, so the perceived delay collapses to the first-byte time rather than the full render.

First byte VS total

There are two clocks. First-byte latency is how long until playback can start. Total render time is how long until the whole clip is ready. For a real-time app, first-byte is what the user feels. For pre-rendered content, total time is what matters. Streaming optimizes the first; the batch endpoint, POST /v1/audio/speech, optimizes for a finished file with speech marks, including WAV.

Audio formats

Five formats: mp3 for general playback, ogg and aac for compressed delivery, pcm at 24kHz mono for raw audio, and u-law for telephony systems. Pick with the Accept header. WAV is not a streaming format; use the batch speech endpoint when you need a WAV file.

Browser playback

In the browser, read the response stream and feed chunks to the Web Audio API for gapless playback. Buffer a little to smooth network jitter, but not so much that you add latency. The tradeoff is buffer size against responsiveness.

Telephony

For phone systems, stream u-law at 8kHz, the format telephony expects. Paired with SIP or a voice agent, streaming TTS puts generated speech straight onto a call with minimal delay.

Live app patterns

Voice agents, live captioning, live translation, in-game dialogue, read-aloud buttons and spoken alerts all depend on audio that starts fast and keeps flowing. A pause before speech reads as a fault in every one of them; an audiobook pipeline, by contrast, cares about total render time because nobody is waiting live. Match the endpoint to the use.

Code samples

curl -X POST https://api.speechify.ai/v1/audio/stream \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Accept: audio/mpeg" \
  -H "Content-Type: application/json" \
  -d '{"input": "Playing as it streams.", "voice_id": "sabrina", "model": "simba-3.2"}' \
  --output speech.mp3
FAQ

Frequently asked questions

What is streaming text-to-speech?
Streaming text-to-speech sends audio back in chunks over a single HTTP connection as it is generated, instead of returning one finished file. Your app decodes and plays the bytes as they arrive, so the first sound reaches the listener after the first byte rather than after the whole clip renders.
How fast is the first byte?
On Simba 3.2 the streaming endpoint returns its first byte in 56 ms at the median and 102 ms at the 90th percentile, measured on the production US East path on 15 Sep 2026. The site rounds that to <100ms. It is a measured median, not a guarantee. Independent benchmarks time the same model to the first audible sound from their own runners: Coval measured a 106 ms median over the 24 hours to 24 Sep 2026, and Voice Arena read 123 ms p50 on its US English board on 24 Sep 2026. Those include their network and any leading silence, so they sit above the first byte.
Is streaming the same as real-time text-to-speech?
Real-time text-to-speech is the use: speech fast enough for someone waiting live. Streaming is how it is delivered: chunked audio that starts playing before generation finishes. Both Simba 3.2 for English and Simba 3.0 for multilingual speech support streaming.
Which formats does streaming support?
mp3 (audio/mpeg), ogg, aac, pcm at 24kHz mono, and u-law for telephony. WAV is not available on streaming; use the batch speech endpoint for WAV. You choose the format with the Accept header.
How much text can I stream per request?
Up to 20,000 characters per request on the streaming endpoint. For longer content, chunk the text and stream the pieces in sequence to keep playback continuous.

Start building

Streaming text-to-speech that plays as it generates, in the browser or over the phone, with a first byte you can measure.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.