GitHubLog in
Response format

Response format

Every endpoint answers with the same JSON envelope. Check ok, branch on data.kind, and keep request_id for support.

  • 200 result
  • 202 job
  • 4xx / 5xx error
Envelope

Four top-level fields

200 OKapplication/json
{
  "ok": true,
  "request_id": "req_8f2c…",
  "data": {
    "kind": "transcript",
    "video_id": "dQw4w9WgXcQ",
    "platform": "youtube",
    "language": "en",
    "source": "captions",
    "segments": [ … ]
  },
  "usage": { "credits_spent": 1, "balance": 99, "bytes": 14233 }
}
  1. okboolean
    true on a 200 and a 202, false on every error.
  2. request_idstring
    Unique per call, and also sent as the X-Request-Id header. Include it when you contact support.
  3. dataobject
    The result. data.kind names its shape; the kinds are listed below.
  4. usageobject
    Credits charged for this call, your remaining balance and the bytes it used. An error carries none because failures are free, and neither does a 202: a job is charged on delivery.
Transcript shape

Segments or one string

Set timestamps in the request to choose how a single transcript comes back. The other key is absent, not null.

data
{
  "source": "captions",
  "segments": [
    { "start": 0, "duration": 3.5,
      "text": "Welcome back to the channel." },
    { "start": 3.5, "duration": 4.1,
      "text": "Today we're building a transcript search tool." },
    { "start": 7.6, "duration": 4.4,
      "text": "First, we fetch the captions for every video." },
    …
  ]
}
Rendered
  1. 0:00.0Welcome back to the channel.
  2. 0:03.5Today we're building a transcript search tool.
  3. 0:07.6First, we fetch the captions for every video.
  4. 0:12.0Then we index the text by timestamp.
The default. Each cue has a start and a duration in seconds.
Source

Where the text came from

data.source says which path produced the transcript. The response shape is the same either way.

"captions"
Caption trackRead from the video's existing captions, in the same request. 1 credit.
"audio"
AI Fallback TranscriptionOn by default: when a video has no captions, or you send mode: "audio", the audio is transcribed. Charged on delivery: 1 credit per started 5 minutes of audio. Send mode: "captions" to never pay for it. How AI Fallback Transcription works →
Async jobs

When you get a 202

AI Fallback Transcription that can't finish inside the inline wait of about 45 seconds answers 202 with a job instead of a transcript. Poll poll_url, or pass a callback_url and the result is POSTed to you. The finished job carries the same transcript data as a 200 and is charged only on delivery.

Polling a job →
202 Accepted
{
  "ok": true,
  "request_id": "req_…",
  "status": "processing",
  "job_id": "asr_m3k1x9qz4vb2p7",
  "poll_url": "/api/v2/transcripts/jobs/asr_m3k1x9qz4vb2p7",
  "data": {
    "kind": "transcript_job",
    "video_id": "dQw4w9WgXcQ",
    "platform": "youtube"
  }
}
Errors

Same envelope, ok is false

An error replaces data with error and carries no usage, because failures are free. Branch on error.code, not the message. When retry_with is present, it names what to change and send again.

All error codes →
422 Unprocessable Entity
{
  "ok": false,
  "request_id": "req_…",
  "error": {
    "code": "no_captions",
    "number": 4103,
    "message": "No caption track (manual or auto-generated) is available. Send mode \"auto\" to transcribe the audio instead.",
    "docs": "https://transcriptfetch.com/docs/errors/no_captions",
    "retry_with": {
      "mode": "auto"
    }
  }
}
Data kinds

data.kind

Branch on data.kind rather than on the endpoint you called: a batch entry and a single fetch carry the same transcript shape.

kindWhat data holds
transcriptA single transcript: video metadata plus segments or text.
video_listA page of videos from a channel, profile, playlist or search, plus next_cursor.
playlist_listA page of playlists: a YouTube channel's playlists or podcasts tab, or a playlist search, plus next_cursor.
channel_listA page of channels from a YouTube channel search, plus next_cursor.
transcript_batchOne entry per requested video in data.results, each with an outcome.
transcript_jobThe 202 envelope and job polling for AI Fallback Transcription.
monitorA monitor. The other monitor endpoints answer monitor_list (your monitors and limits), monitor_event_list (a page of events), monitor_check (a check's result) and monitor_deleted.

Single transcript

The /video endpoint returns kind: "transcript" with either a joined text string or a timestamped segments array, never both. The timestamps request field chooses: true (the default) returns segments, false returns text. The unwanted key is absent, not null, so check for presence, not truthiness. Each segment's start and duration are in seconds, everything needed to render timestamps or export SRT/VTT. source says where the words came from: "captions" for an existing caption track, "audio" for AI Fallback Transcription (on by default when there are no captions; billed by length, see usage.credits_spent).

Every transcript also carries the same metadata block, whichever platform or path served it: video_id (the item's id on its platform: the 11-character YouTube id, TikTok's numeric id, the Instagram shortcode, or the URL for a direct file), url (the canonical link, accepted as-is by every transcript endpoint), platform, title (the post caption on TikTok and Instagram), channel (the creator, as the platform names them: a channel name, a TikTok @handle, an Instagram username), duration in seconds, language (the caption track's code, or the language detected during AI Fallback Transcription), thumbnail_url (TikTok and Instagram serve signed, expiring poster URLs, so copy the image rather than hotlinking it), and source. Every key is always present; null means unknown, never omitted. A cache hit, an inline audio result, a job poll and a webhook delivery all carry exactly these keys.

JSON
{
  "ok": true,
  "request_id": "req_…",
  "data": {
    "kind": "transcript",
    "video_id": "dQw4w9WgXcQ",
    "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "platform": "youtube",
    "title": "Example video",
    "channel": "Example Channel",
    "duration": 212,
    "language": "en",
    "thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg",
    "source": "captions",
    "segments": [
      {
        "start": 0,
        "duration": 3.5,
        "text": "We're no strangers to love"
      }
    ]
  },
  "usage": {
    "credits_spent": 1,
    "balance": 99,
    "bytes": 14233
  }
}

The /channel, /playlist, and /search endpoints return kind: "video_list", a videos array of metadata plus a next_cursor for pagination. Fetch the transcripts separately via the batch endpoint. A channel, profile or playlist that does not exist (or is not public) answers 404 with error.code: "not_found"; an existing one with nothing to list answers 200 with an empty videos array, free of charge. An upstream fault while listing answers 503, never an empty page. A list can also carry a notice saying why a page is empty or may be incomplete, such as "This channel has no Live tab."

JSON
{
  "ok": true,
  "request_id": "req_…",
  "data": {
    "kind": "video_list",
    "source": "search",
    "platform": "tiktok",
    "videos": [
      {
        "videoId": "7398765432101234567",
        "url": "https://www.tiktok.com/@examplecreator/video/7398765432101234567",
        "title": "the post caption",
        "duration": 58,
        "channel": "@examplecreator",
        "publishedAt": "2026-08-30T14:02:00Z",
        "stats": {
          "plays": 1240000
        }
      }
    ],
    "next_cursor": "eyJvIjoxMH0"
  },
  "usage": {
    "credits_spent": 1,
    "balance": 98,
    "bytes": 0
  }
}

Playlist and channel lists

Some YouTube listings return playlists or channels instead of videos. A channel's playlists or podcasts tab, and a search with type: "playlist", answer with kind: "playlist_list" and a playlists array; each row carries playlistId, url, title, channel and videoCount. A search with type: "channel" answers with kind: "channel_list" and a channels array; each row carries channelId, url, handle, title, subscribers (read from YouTube's rounded label), videoCount (always null, because YouTube's channel results no longer show one) and description.

Every row's url is accepted as-is by the playlist or channel endpoint, and both kinds page with next_cursor exactly like a video list. The channel and search references show a full example of each.

Batch

The batch endpoint returns kind: "transcript_batch" with one entry per requested video in data.results. Batch entries carry both text and segments (the timestamps flag is not accepted on batch), plus outcome: ok, processing, or error. A failed entry carries the same error block as a request-level failure; entries without text do not ship the text fields at all.

Entries with no caption track go to AI Fallback Transcription by default: those come back with outcome: "processing" plus a job_id and poll_url, cost nothing on the batch call, and are billed on delivery at the audio rate. Re-send the same batch once they finish and the text is returned normally; polling is optional. Send mode: "captions" to have captionless entries fail as error with code no_captions (and a retry_with naming the audio mode) instead.

JSON
{
  "ok": true,
  "request_id": "req_…",
  "data": {
    "kind": "transcript_batch",
    "results": [
      {
        "video_id": "dQw4w9WgXcQ",
        "outcome": "ok",
        "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "title": "Example video",
        "channel": "Example Channel",
        "duration": 212,
        "language": "en",
        "thumbnail_url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/mqdefault.jpg",
        "source": "captions",
        "text": "Full transcript text …",
        "segments": [
          {
            "start": 0,
            "duration": 3.5,
            "text": "Full transcript …"
          }
        ],
        "bytes": 14233
      },
      {
        "video_id": "9bZkp7q19f0",
        "outcome": "processing",
        "job_id": "asr_…",
        "poll_url": "/api/v2/transcripts/jobs/asr_…"
      },
      {
        "video_id": "jNQXAC9IVRw",
        "outcome": "error",
        "error": {
          "code": "no_captions",
          "number": 4103,
          "message": "No caption track (manual or auto-generated) is available.",
          "docs": "https://transcriptfetch.com/docs/errors/no_captions",
          "retry_with": {
            "mode": "auto"
          }
        }
      }
    ]
  },
  "usage": {
    "credits_spent": 1,
    "balance": 97
  }
}

Field reference

  • data.kind: transcript, video_list, playlist_list, channel_list, transcript_batch, transcript_job, or me; the monitor endpoints answer with monitor, monitor_list, monitor_event_list, monitor_check and monitor_deleted.
  • text / segments: exactly one on single-transcript responses (chosen by timestamps); both on batch entries and completed job results. Segments are { start, duration, text } cues in seconds (speaker added when diarized).
  • next_cursor: pass back as cursor for the next page of a list, null when exhausted.
  • outcome (batch results): ok, processing, or error; failed entries carry error (the standard block), and processing entries (captionless videos in AI Fallback Transcription) carry job_id and poll_url.
  • video_id, url, platform, title, channel, duration, language, thumbnail_url (transcripts): the metadata block above, on every transcript-bearing response including job polls, webhooks and batch entries. Always present; null when unknown.
  • source (transcripts): "captions" or "audio".
  • usage.credits_spent: credits charged for this request. Failures are never charged and carry no usage block.
  • usage.balance: remaining credit balance (null for unlimited accounts).
  • usage.bytes: upstream bytes this request consumed (absent on batch, which reports per-entry bytes).