How a request resolves
- Caption track foundManual or auto-generated captions are read and returned.200
- No captions, under 20 minutesAI Fallback Transcription runs while the request is held open for up to 45 seconds. No polling.200
- Longer, or still running after 45 sAnswers with a job to poll, or POSTs the result to your callback_url. Retrying is safe and hits the cache once it is done.202
Body parameters
A video URL or 11-character YouTube video ID. Accepts YouTube (watch, youtu.be, /shorts/), TikTok, and Instagram URLs, plus direct media file URLs (mp4/mp3/wav/…). Under the default mode, a video without captions is transcribed by AI Fallback Transcription (see mode).
Where the text may come from. "auto" (the default) reads the caption track, and when there is none transcribes the audio with AI Fallback Transcription, charged on delivery at 1 credit per started 5 minutes of audio, minimum 1 (a 20-minute video is 4; the 4-hour cap is 48). "captions" reads an existing caption track only and answers 422 no_captions when there is none: the one mode that never charges for audio. "audio" skips captions and always uses AI Fallback Transcription, at the same rate. A caption fetch is always 1 credit. AI Fallback Transcription on short media may finish inline after a wait of up to 45 seconds; longer work returns 202 with a job to poll or deliver by callback.
Which form the transcript comes back in. true (the default) returns the `segments` array, each with start, duration and text. false returns a single joined `text` string instead. On the single-transcript 200 exactly one of the two is present, never both, since segments already contain every word the joined text does; job results and batch entries carry both text and segments. The older strings "segment" and "none" mean the same two things and are still accepted.
Where to POST the finished transcript when a request escalates to AI Fallback Transcription, instead of polling the job. The delivery body is a trimmed envelope - {ok, status, job_id, data} on success, {ok, status, job_id, error} on failure - without the usage and request_id the poll URL adds. When a signing secret is configured on our side, the body is signed with HMAC-SHA256 over the exact bytes and sent as an X-TranscriptFetch-Signature: sha256=<hex> header so you can verify it came from us. Must be a public https URL on the standard port; the URL is checked again at delivery time, so an unreachable or private address still gets a 202 but never receives a delivery. The job stays pollable either way, so a missed delivery is never a lost transcript.
Legacy alias for "mode", still supported. true is identical to "mode": "audio" (always AI Fallback Transcription); omitted or false is "mode": "auto", the default, which still uses AI Fallback Transcription when a video has no captions. Send one or the other, not both. false never turned the fallback off, which is why the field was replaced: send "mode": "captions" for that.
{
"ok": true,
"request_id": "req_…",
"data": {
"kind": "transcript",
"video_id": "dQw4w9WgXcQ",
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"platform": "youtube",
"title": "Example video",
"channel": "Example Channel",
"duration": 212,
"language": "en",
"thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg",
"source": "captions",
"segments": [
{
"start": 0,
"duration": 3.5,
"text": "We're no strangers to love"
}
]
},
"usage": {
"credits_spent": 1,
"balance": 99,
"bytes": 14233
}
}Every YouTube endpoint
- POSTTranscriptFetch a YouTube video's transcript over REST: the request, every parameter, the JSON response with timestamped segments, and curl, Python and Node examples.
- POSTSearchSearch YouTube over REST without a Data API key or quota: the request, the upload date, duration, captions and sort filters, a real response, and curl, Python and Node examples.
- POSTChannel VideosList a YouTube channel's videos over REST without a Data API key: the request, the tab, sort and query options, since_video_id polling, a real response, and curl, Python and Node examples.
- POSTPlaylistList a YouTube playlist's videos over REST without a Data API key: the request with a playlist URL or PL id, cursor paging, a real response, and curl, Python and Node examples.
- POSTShortsList and transcribe a channel's YouTube Shorts over REST: the channel endpoint with tab shorts, a real Shorts listing, how a Short's transcript is fetched, and curl, Python and Node examples.
- POSTChannel MonitorWatch a YouTube channel for new videos over REST: create a monitor, receive each upload by signed webhook or from the events list, optionally with its transcript, with curl, Python and Node examples.
Same call, other platforms:TikTok Transcript APIInstagram Transcript API
Frequently asked questions
Which YouTube URLs does the transcript endpoint accept?
Standard watch URLs, youtu.be short links, Shorts, embed and live URLs, and the bare video id. The full list, with an example of each, is on the YouTube URL formats page.
What happens when a YouTube video has no captions?
With the default mode, auto, AI Fallback Transcription transcribes the audio instead. Media under 20 minutes is held open for up to 45 seconds and returns the finished transcript inline; longer media, or a transcription still running when the hold expires, answers 202 with a job_id and poll_url. Send mode captions to fail with no_captions instead.
Do I get timestamps?
Yes by default: data.segments carries start, duration and text for every caption cue. Send timestamps false to get a single data.text string instead; a response never carries both.
How is a YouTube transcript billed?
A transcript served from captions costs 1 credit. One from AI Fallback Transcription is billed by duration, on delivery. A request that fails, or a video that is private, removed or otherwise unavailable, costs nothing.