There are two practical ways to fetch YouTube transcripts in Python: read available captions with an open-source library, or use a hosted API. This guide shows both, including where a local script differs from a production integration.
We build TranscriptFetch’s YouTube Transcript API, which handles transcript retrieval, audio fallback, and video discovery through one service. The examples below use its official Python SDK alongside the separate open-source youtube-transcript-api library, so you can choose the approach that fits your application.
TL;DR: For Python applications, we recommend TranscriptFetch’s YouTube Transcript API and official Python SDK to fetch transcripts, handle audio fallback, and collect videos through batches and channel listings. Handle pending audio jobs and per-video failures explicitly. The open-source youtube-transcript-api library remains a useful option for small local scripts that only need available captions.
Option 1: the free youtube-transcript-api library
The youtube-transcript-api library retrieves existing YouTube captions without a TranscriptFetch account or API key. This example uses version 1.2.4. Start with a virtual environment on a supported Python release:
python -m venv .venv
source .venv/bin/activate
python -m pip install youtube-transcript-api==1.2.4On Windows, activate the environment with .venv\Scripts\Activate.ps1 in PowerShell instead. Save this as fetch_oss.py and run python fetch_oss.py:
from youtube_transcript_api import YouTubeTranscriptApi
api = YouTubeTranscriptApi()
fetched = api.fetch("aircAruvnKk", languages=["en", "es"])
print(" ".join(snippet.text for snippet in fetched))
for snippet in fetched:
print(snippet.start, snippet.duration, snippet.text)
# Use dictionaries when saving JSON or integrating with existing code.
rows = fetched.to_raw_data()
print(fetched.language_code, fetched.is_generated)Pass a video ID, not the full URL. The language list is a preference order for available tracks, not a request to translate the video. fetch() returns an iterable FetchedTranscript object; to_raw_data() converts its snippets to dictionaries.
Older tutorials use the removed static method YouTubeTranscriptApi.get_transcript(). For the pinned version above, instantiate the client and call fetch() instead, including when you supply languages.
Where the library needs extra work
The library reads caption tracks; it does not transcribe a video’s audio when a track is missing. It exposes transcript metadata such as language and whether a track is generated, but it is not a video discovery or channel listing service.
Its maintainer documents RequestBlocked and IpBlocked failures, particularly from cloud-hosted IP ranges. That does not mean every deployment fails, but a successful laptop test is not enough to establish production reliability. Compare the same request from your actual host. Our production troubleshooting guide covers that diagnosis.
For many videos, you also need to manage concurrency, application-level retries, and saved results. A missing language track or inaccessible video is not necessarily a transient error; retrying every failure identically wastes work.
Option 2: TranscriptFetch’s YouTube Transcript API
TranscriptFetch’s YouTube Transcript API moves the retrieval infrastructure into a hosted service. Its Python SDK provides single-video retrieval, batch results, and channel, playlist, and search listing. You still handle API errors and pending work, but you do not need to operate the underlying scraper or proxy pool.
These examples target the published transcriptfetch-sdk 2.3.0 package and the v2 API. Install it and set a key from your dashboard:
python -m pip install transcriptfetch-sdk==2.3.0
export TRANSCRIPTFETCH_API_KEY='tf_live_YOUR_KEY'In PowerShell, use $env:TRANSCRIPTFETCH_API_KEY='tf_live_YOUR_KEY'. The SDK reads this variable automatically. Keep the key server-side and out of source control. The Python SDK docs describe the client, and its source repository records the implementation.
Fetch a video and handle audio jobs
By default, the API tries captions and can fall back to audio transcription. Some audio work finishes inline; other requests return HTTP 202 with a job ID. That response means processing has started, not that transcript text is ready.
Save this as fetch_api.py and run python fetch_api.py:
import time
from transcriptfetch import TranscriptFetch, APIError
def wait_for_transcript(tf, result, max_wait_seconds=120):
deadline = time.monotonic() + max_wait_seconds
while result.status == "processing":
if not result.job_id:
raise RuntimeError("Pending response has no job ID")
if time.monotonic() >= deadline:
raise TimeoutError(f"Resume polling job {result.job_id}")
time.sleep(2)
# SDK 2.3.0 raises APIError for a failed job envelope.
result = tf.transcripts.job(result.job_id)
if result.status == "failed":
message = result.error.message if result.error else "Transcription failed"
raise RuntimeError(message)
return result
def main():
# api_key falls back to TRANSCRIPTFETCH_API_KEY.
with TranscriptFetch(timeout=60, max_retries=2) as tf:
result = wait_for_transcript(
tf, tf.transcripts.video("https://youtu.be/aircAruvnKk")
)
text = result.text
if text is None:
text = " ".join(segment.text for segment in result.segments)
print(result.title)
print(text)
for segment in result.segments[:3]:
print(segment.start, segment.duration, segment.text)
if __name__ == "__main__":
try:
main()
except APIError as error:
print(error.status, error.code, error.request_id)
raiseThe single-video HTTP response contains either segments or text, selected by timestamps. The default returns segments; SDK 2.3.0 represents the absent text field as None. Joining segment text locally gives you plain text without another request. Timing values are seconds. Metadata can be None when the service cannot determine it.
The polling deadline limits how long this script keeps checking. It does not cancel server-side transcription, so retain the job ID to resume later. SDK 2.3.0 raises an exception when polling returns a failed job envelope; do not treat failed jobs as completed transcripts.
Request captions only or plain text
The API supports mode, timestamps, and callback_url, but Python SDK 2.3.0 does not expose them as arguments to video(). Use its generic _request() method for these fields. That method returns the raw envelope rather than a normalized transcript model.
# plain_api.py
from transcriptfetch import TranscriptFetch
with TranscriptFetch(timeout=60) as tf:
envelope = tf._request(
"POST",
"/api/v2/transcripts/video",
body={
"video": "aircAruvnKk",
"mode": "captions",
"timestamps": False,
},
idempotent=True,
)
print(envelope["data"]["text"])This captions-only request either returns finished text or raises an API error; it does not start audio transcription. Switching to auto or audio requires handling a possible job response. The endpoint has no language-selection parameter; the response’s language reports the retrieved or detected language.
The underlying endpoint is POST https://transcriptfetch.com/api/v2/transcripts/video, with the input in video. Raw responses wrap results in data and include usage information. See the endpoint reference for the complete contract.
Fetch a batch and inspect each outcome
The batch limit is 50 entries on free, Basic, and Pro plans, and 500 on Mega and Scale. Here we explicitly choose captions-only mode so videos without a usable caption track return an individual error instead of starting audio work.
# batch_api.py
from transcriptfetch import TranscriptFetch
with TranscriptFetch(timeout=120) as tf:
batch = tf.transcripts.batch(
["aircAruvnKk", "dQw4w9WgXcQ"], mode="captions"
)
for item in batch.results:
if item.outcome == "ok":
print(item.video_id, item.text)
elif item.outcome == "processing":
print(item.video_id, "Pending job:", item.job_id)
else:
code = item.error.code if item.error else "unknown"
print(item.video_id, "Error:", code)Batch outcomes are ok, processing, or error. The processing branch matters if you change the mode to auto: a captionless entry may return a job ID. Poll that job, or re-send the batch after processing finishes. A missing-caption error is a code inside the error object, not an outcome named no_transcript.
Successful batch entries can carry both text and segments. Count entries with outcome == "ok" when reporting delivered transcripts; the length of the results array also includes failures and pending jobs. Authentication, credit, and other request-level errors still raise exceptions for the whole call.
Discover channel videos before fetching transcripts
TranscriptFetch’s channel, playlist, and search endpoints return video metadata. Fetch the transcript separately for each result, or send the results to the batch endpoint. You do not need a Google API key for these hosted listing calls.
This example requests ten videos per page and fetches captions for up to 20 videos. The cap keeps a first run small:
# channel_api.py
from itertools import islice
from transcriptfetch import TranscriptFetch
with TranscriptFetch(timeout=120) as tf:
videos = list(islice(
tf.transcripts.iter_channel("@GoogleDevelopers", limit=10), 20
))
if videos:
inputs = [video.url or video.video_id for video in videos]
batch = tf.transcripts.batch(inputs, mode="captions")
delivered = [item for item in batch.results if item.outcome == "ok"]
print(f"Delivered {len(delivered)} of {len(inputs)} transcripts")
for item in delivered:
print(item.video_id, item.text)iter_channel() follows cursors for you. For larger imports, process bounded batches rather than collecting an entire channel in memory, and persist your progress. iter_playlist() and iter_search() support the same discovery pattern. The whole-channel guide explains the collection workflow; pagination docs describe cursors.
Credits, errors, and choosing a workflow
Caption retrieval costs one credit per successful transcript. Audio transcription costs one credit per started five minutes, with a minimum of one, charged on delivery. Non-empty listing pages have their own one-credit charge, and job polling is free. Failed requests are free, but a captionless video that is successfully transcribed from audio is not. Consult billing for the full rules.
The SDK retries HTTP 429 and 5xx responses with backoff and honors Retry-After. It preserves the idempotency key across those HTTP retries. Connection errors, timeouts, authentication problems, and exhausted credits still need application handling; inspect the exception rather than retrying everything indefinitely.
| Requirement | Open-source caption library | TranscriptFetch |
|---|---|---|
| Small local caption script | Often a good fit | Optional hosted alternative |
| Transcript language preferences | Ordered available-track selection | No language-selection request field |
| Video discovery | Separate implementation | Channel, playlist, and search listing |
| Captionless video | Separate speech-to-text system | Audio fallback or explicit captions-only mode |
| Deployment infrastructure | You manage upstream access | Retrieval infrastructure is hosted |
| Batch results | Your own loop and bookkeeping | Per-video outcomes from a batch endpoint |
Use the library when a small caption script fits the job. When you need hosted retrieval for an application, TranscriptFetch’s YouTube Transcript API is the product demonstrated here. Start with its Python SDK docs, or use the YouTube MCP Server to make retrieval available to an AI client. The Node.js guide covers the corresponding JavaScript integration.
