Four top-level fields
{
"ok": true,
"request_id": "req_8f2c…",
"data": {
"kind": "transcript",
"video_id": "dQw4w9WgXcQ",
"platform": "youtube",
"language": "en",
"source": "captions",
"segments": [ … ]
},
"usage": { "credits_spent": 1, "balance": 99, "bytes": 14233 }
}- true on a 200 and a 202, false on every error.
okboolean - Unique per call, and also sent as the X-Request-Id header. Include it when you contact support.
request_idstring - The result. data.kind names its shape; the kinds are listed below.
dataobject - Credits charged for this call, your remaining balance and the bytes it used. An error carries none because failures are free, and neither does a 202: a job is charged on delivery.
usageobject
Segments or one string
Set timestamps in the request to choose how a single transcript comes back. The other key is absent, not null.
{
"source": "captions",
"segments": [
{ "start": 0, "duration": 3.5,
"text": "Welcome back to the channel." },
{ "start": 3.5, "duration": 4.1,
"text": "Today we're building a transcript search tool." },
{ "start": 7.6, "duration": 4.4,
"text": "First, we fetch the captions for every video." },
…
]
}- 0:00.0Welcome back to the channel.
- 0:03.5Today we're building a transcript search tool.
- 0:07.6First, we fetch the captions for every video.
- 0:12.0Then we index the text by timestamp.
Where the text came from
data.source says which path produced the transcript. The response shape is the same either way.
"captions""audio"mode: "audio", the audio is transcribed. Charged on delivery: 1 credit per started 5 minutes of audio. Send mode: "captions" to never pay for it. How AI Fallback Transcription works →When you get a 202
AI Fallback Transcription that can't finish inside the inline wait of about 45 seconds answers 202 with a job instead of a transcript. Poll poll_url, or pass a callback_url and the result is POSTed to you. The finished job carries the same transcript data as a 200 and is charged only on delivery.
{
"ok": true,
"request_id": "req_…",
"status": "processing",
"job_id": "asr_m3k1x9qz4vb2p7",
"poll_url": "/api/v2/transcripts/jobs/asr_m3k1x9qz4vb2p7",
"data": {
"kind": "transcript_job",
"video_id": "dQw4w9WgXcQ",
"platform": "youtube"
}
}Same envelope, ok is false
An error replaces data with error and carries no usage, because failures are free. Branch on error.code, not the message. When retry_with is present, it names what to change and send again.
{
"ok": false,
"request_id": "req_…",
"error": {
"code": "no_captions",
"number": 4103,
"message": "No caption track (manual or auto-generated) is available. Send mode \"auto\" to transcribe the audio instead.",
"docs": "https://transcriptfetch.com/docs/errors/no_captions",
"retry_with": {
"mode": "auto"
}
}
}data.kind
Branch on data.kind rather than on the endpoint you called: a batch entry and a single fetch carry the same transcript shape.
| kind | What data holds |
|---|---|
| transcript | A single transcript: video metadata plus segments or text. |
| video_list | A page of videos from a channel, profile, playlist or search, plus next_cursor. |
| playlist_list | A page of playlists: a YouTube channel's playlists or podcasts tab, or a playlist search, plus next_cursor. |
| channel_list | A page of channels from a YouTube channel search, plus next_cursor. |
| transcript_batch | One entry per requested video in data.results, each with an outcome. |
| transcript_job | The 202 envelope and job polling for AI Fallback Transcription. |
| monitor | A monitor. The other monitor endpoints answer monitor_list (your monitors and limits), monitor_event_list (a page of events), monitor_check (a check's result) and monitor_deleted. |
Single transcript
The /video endpoint returns kind: "transcript" with either a joined text string or a timestamped segments array, never both. The timestamps request field chooses: true (the default) returns segments, false returns text. The unwanted key is absent, not null, so check for presence, not truthiness. Each segment's start and duration are in seconds, everything needed to render timestamps or export SRT/VTT. source says where the words came from: "captions" for an existing caption track, "audio" for AI Fallback Transcription (on by default when there are no captions; billed by length, see usage.credits_spent).
Every transcript also carries the same metadata block, whichever platform or path served it: video_id (the item's id on its platform: the 11-character YouTube id, TikTok's numeric id, the Instagram shortcode, or the URL for a direct file), url (the canonical link, accepted as-is by every transcript endpoint), platform, title (the post caption on TikTok and Instagram), channel (the creator, as the platform names them: a channel name, a TikTok @handle, an Instagram username), duration in seconds, language (the caption track's code, or the language detected during AI Fallback Transcription), thumbnail_url (TikTok and Instagram serve signed, expiring poster URLs, so copy the image rather than hotlinking it), and source. Every key is always present; null means unknown, never omitted. A cache hit, an inline audio result, a job poll and a webhook delivery all carry exactly these keys.
{ "ok": true, "request_id": "req_…", "data": { "kind": "transcript", "video_id": "dQw4w9WgXcQ", "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ", "platform": "youtube", "title": "Example video", "channel": "Example Channel", "duration": 212, "language": "en", "thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg", "source": "captions", "segments": [ { "start": 0, "duration": 3.5, "text": "We're no strangers to love" } ] }, "usage": { "credits_spent": 1, "balance": 99, "bytes": 14233 } }
Video list (channel / playlist / search)
The /channel, /playlist, and /search endpoints return kind: "video_list", a videos array of metadata plus a next_cursor for pagination. Fetch the transcripts separately via the batch endpoint. A channel, profile or playlist that does not exist (or is not public) answers 404 with error.code: "not_found"; an existing one with nothing to list answers 200 with an empty videos array, free of charge. An upstream fault while listing answers 503, never an empty page. A list can also carry a notice saying why a page is empty or may be incomplete, such as "This channel has no Live tab."
{ "ok": true, "request_id": "req_…", "data": { "kind": "video_list", "source": "search", "platform": "tiktok", "videos": [ { "videoId": "7398765432101234567", "url": "https://www.tiktok.com/@examplecreator/video/7398765432101234567", "title": "the post caption", "duration": 58, "channel": "@examplecreator", "publishedAt": "2026-08-30T14:02:00Z", "stats": { "plays": 1240000 } } ], "next_cursor": "eyJvIjoxMH0" }, "usage": { "credits_spent": 1, "balance": 98, "bytes": 0 } }
Playlist and channel lists
Some YouTube listings return playlists or channels instead of videos. A channel's playlists or podcasts tab, and a search with type: "playlist", answer with kind: "playlist_list" and a playlists array; each row carries playlistId, url, title, channel and videoCount. A search with type: "channel" answers with kind: "channel_list" and a channels array; each row carries channelId, url, handle, title, subscribers (read from YouTube's rounded label), videoCount (always null, because YouTube's channel results no longer show one) and description.
Every row's url is accepted as-is by the playlist or channel endpoint, and both kinds page with next_cursor exactly like a video list. The channel and search references show a full example of each.
Batch
The batch endpoint returns kind: "transcript_batch" with one entry per requested video in data.results. Batch entries carry both text and segments (the timestamps flag is not accepted on batch), plus outcome: ok, processing, or error. A failed entry carries the same error block as a request-level failure; entries without text do not ship the text fields at all.
Entries with no caption track go to AI Fallback Transcription by default: those come back with outcome: "processing" plus a job_id and poll_url, cost nothing on the batch call, and are billed on delivery at the audio rate. Re-send the same batch once they finish and the text is returned normally; polling is optional. Send mode: "captions" to have captionless entries fail as error with code no_captions (and a retry_with naming the audio mode) instead.
{ "ok": true, "request_id": "req_…", "data": { "kind": "transcript_batch", "results": [ { "video_id": "dQw4w9WgXcQ", "outcome": "ok", "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ", "title": "Example video", "channel": "Example Channel", "duration": 212, "language": "en", "thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg", "source": "captions", "text": "Full transcript text …", "segments": [ { "start": 0, "duration": 3.5, "text": "Full transcript …" } ], "bytes": 14233 }, { "video_id": "9bZkp7q19f0", "outcome": "processing", "job_id": "asr_…", "poll_url": "/api/v2/transcripts/jobs/asr_…" }, { "video_id": "jNQXAC9IVRw", "outcome": "error", "error": { "code": "no_captions", "number": 4103, "message": "No caption track (manual or auto-generated) is available.", "docs": "https://transcriptfetch.com/docs/errors/no_captions", "retry_with": { "mode": "auto" } } } ] }, "usage": { "credits_spent": 1, "balance": 97 } }
Field reference
data.kind:transcript,video_list,playlist_list,channel_list,transcript_batch,transcript_job, orme; the monitor endpoints answer withmonitor,monitor_list,monitor_event_list,monitor_checkandmonitor_deleted.text/segments: exactly one on single-transcript responses (chosen bytimestamps); both on batch entries and completed job results. Segments are{ start, duration, text }cues in seconds (speakeradded when diarized).next_cursor: pass back ascursorfor the next page of a list,nullwhen exhausted.outcome(batch results):ok,processing, orerror; failed entries carryerror(the standard block), andprocessingentries (captionless videos in AI Fallback Transcription) carryjob_idandpoll_url.video_id,url,platform,title,channel,duration,language,thumbnail_url(transcripts): the metadata block above, on every transcript-bearing response including job polls, webhooks and batch entries. Always present;nullwhen unknown.source(transcripts):"captions"or"audio".usage.credits_spent: credits charged for this request. Failures are never charged and carry nousageblock.usage.balance: remaining credit balance (nullfor unlimited accounts).usage.bytes: upstream bytes this request consumed (absent on batch, which reports per-entrybytes).