Streaming text-to-speech
Streaming text-to-speech that plays as it generates, in the browser or over the phone, with a first byte you can measure.
Definition
Streaming text-to-speech returns audio in chunks as it is generated, over a single HTTP connection. Your app decodes and plays the bytes as they arrive, so the listener hears the first sound almost immediately instead of waiting for the whole clip to render. Real-time text-to-speech is what that makes possible: speech fast enough for a live app, where someone is waiting to hear the answer.
Measured latency
SpeechifyAI’s streaming endpoint returns its first byte in 56 ms at the median and 102 ms at the 90th percentile on Simba 3.2, measured on the production US East path on 15 Sep 2026. The site rounds that to <100ms; the number is a measured median, not a guarantee.
Two independent benchmarks time Simba 3.2 from their own runners to the first sound: Coval measured a 106 ms median time to first audio over the 24 hours to 24 Sep 2026, and Voice Arena read 123 ms p50 on its US English board the same day. Their readings include the network to each runner and any silence before the first sound, so they read higher than our own first byte, and neither is its source. Our reviewed snapshots: Coval, Voice Arena.
How streaming works
You call POST /v1/audio/stream. The response is chunked HTTP: audio bytes flow back continuously. Your client reads the stream, decodes each chunk, and feeds it to the audio output. Playback starts on the first chunk and continues as generation continues, so the perceived delay collapses to the first-byte time rather than the full render.
First byte VS total
There are two clocks. First-byte latency is how long until playback can start. Total render time is how long until the whole clip is ready. For a real-time app, first-byte is what the user feels. For pre-rendered content, total time is what matters. Streaming optimizes the first; the batch endpoint, POST /v1/audio/speech, optimizes for a finished file with speech marks, including WAV.
Audio formats
Five formats: mp3 for general playback, ogg and aac for compressed delivery, pcm at 24kHz mono for raw audio, and u-law for telephony systems. Pick with the Accept header. WAV is not a streaming format; use the batch speech endpoint when you need a WAV file.
Browser playback
In the browser, read the response stream and feed chunks to the Web Audio API for gapless playback. Buffer a little to smooth network jitter, but not so much that you add latency. The tradeoff is buffer size against responsiveness.
Telephony
For phone systems, stream u-law at 8kHz, the format telephony expects. Paired with SIP or a voice agent, streaming TTS puts generated speech straight onto a call with minimal delay.
Live app patterns
Voice agents, live captioning, live translation, in-game dialogue, read-aloud buttons and spoken alerts all depend on audio that starts fast and keeps flowing. A pause before speech reads as a fault in every one of them; an audiobook pipeline, by contrast, cares about total render time because nobody is waiting live. Match the endpoint to the use.
Code samples
curl -X POST https://api.speechify.ai/v1/audio/stream \
-H "Authorization: Bearer $SPEECHIFY_API_KEY" \
-H "Accept: audio/mpeg" \
-H "Content-Type: application/json" \
-d '{"input": "Playing as it streams.", "voice_id": "sabrina", "model": "simba-3.2"}' \
--output speech.mp3Frequently asked questions
What is streaming text-to-speech?
How fast is the first byte?
Is streaming the same as real-time text-to-speech?
Which formats does streaming support?
How much text can I stream per request?
Continue exploring speech.
Start building
Streaming text-to-speech that plays as it generates, in the browser or over the phone, with a first byte you can measure.