GitHubLog in
Limits & billing

Rate limits

Your plan's request allowance is shared across your API keys, dashboard, and MCP connections. Use the response headers to pace requests and retry when capacity becomes available.

Limits by plan

PlanRequests per minuteConcurrent extraction slots
Free102
Basic602
Pro3003
Mega90010
Scale900; custom increases available10; custom increases available

The request allowance is shared across all API keys, dashboard sessions, and MCP connections belonging to your account. Creating another key does not add capacity. Requests are counted over a sliding 60-second window. A lower limit set on an individual key also applies; contact support for a custom increase.

Response headers

Rate-limited endpoints report the current allowance in these headers. This example shows a Basic account:

HTTP
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 58
X-RateLimit-Reset: 1717619400

X-RateLimit-Reset is a Unix timestamp in seconds. When a key has a lower limit, the headers describe whichever allowance is more restrictive.

Concurrent work and batches

An individual extraction request occupies one slot until it returns. A batch counts as one request against the minute allowance, but reserves a slot for each of its workers: up to two on Free and Basic, three on Pro, and four on Mega and Scale. Batch items run through that pool, so a large batch does not start every video at once. Concurrent batches and individual requests share the same account capacity.

If the account is busy, wait before submitting more work. A batch may contain successful items alongside retryable errors; submit only the failed items in a new request after the indicated delay, using a new Idempotency-Key if you use one. Successful deliveries are billed normally; rejected items are free.

Job-status polling

GET /transcripts/jobs/{jobId} uses a separate allowance: twice your account's request limit, with a minimum of 120 polls per minute. This allowance is shared across your keys and connections. The MCP get_credits tool uses this lightweight allowance too. Polls do not consume extraction slots or credits. Repeating a transcript request is a new request, not a job-status poll.

Handling 429 responses

When you exceed the limit you get 429 Too Many Requests. The wait time comes from the Retry-After header (seconds), and the body is the standard error envelope:

HTTP
HTTP/1.1 429 Too Many Requests
Retry-After: 12

{
  "ok": false,
  "request_id": "req_8f2c1a90b3d44e01",
  "error": {
    "code": "rate_limited",
    "number": 2101,
    "message": "Too many requests. Slow down and retry.",
    "docs": "https://transcriptfetch.com/docs/errors/rate_limited"
  }
}

A 429 never consumes credits.

A 429 means your account or key has reached a request or work-capacity limit. When an upstream platform blocks or fails a fetch, the API answers 503 with code upstream_error (5001) instead. If request admission is temporarily unavailable, requests are rejected before extraction with 503 and Retry-After.

Honor Retry-After. When you hit a 429, wait the number of seconds in the Retry-After header before retrying rather than hammering the endpoint immediately.