On this page
The open-source youtube-transcript-api library is the first thing most people reach for, and it is good. It is free, needs no key, and on your own machine it usually just works. Then you deploy it and it starts raising RequestBlocked or IpBlocked.
Why it fails on a server #
The library talks to YouTube directly, so YouTube sees the address your code runs from. The project's README says what happens next:
YouTube has started blocking most IPs that are known to belong to cloud providers (like AWS, Google Cloud Platform, Azure, etc.)
Its fix is rotating residential proxies, and it is honest about them: "using a proxy doesn't guarantee that you won't be blocked".
The official YouTube Data API doesn't help with other people's videos. Downloading a caption track requires permission to edit the video.
Your options #
| Option | You run | Good for |
|---|---|---|
| The library on your own machine | Nothing extra | Scripts, research, one-off jobs |
| The library with rotating residential proxies | A proxy account and retry logic | A pipeline you want fully in your hands |
| A hosted transcript API | Nothing | Servers, automations, products |
If the job runs on your laptop, keep the library. If it runs on a server, you need your own proxies or a hosted API.
Hosted transcript APIs include Supadata, TranscriptAPI.com, ScrapeCreators, actors on Apify, and CleanScript. We haven't compared them side by side, and their plans change often, so read each pricing page.
What CleanScript does differently #
YouTube's automatic captions have no punctuation, no capitals and misheard names. That makes them hard to read and worse to give a model. CleanScript corrects them, and returns a creator's own captions as written.
We measured the correction on 15 public videos in English, French, Spanish and German. For each, we compared the automatic captions word by word with the creator's own captions, before and after correction, ignoring case and punctuation:
- 407 words fixed: wrong in the automatic captions, right after correction.
- 70 words broken: right before, wrong after.
Correction fixes about six words for each one it breaks, and never changes a number or a negation.
It also transcribes a video without captions instead of failing (up to 15 minutes long), takes TikTok and Instagram links in the same call (in beta), and returns fields with quotes instead of the text when you ask.
What it costs #
New accounts get 100 free credits, no card. A transcript is 1 credit per started 10 minutes of video, and a failed request is free. See pricing for packs.
What it doesn't do #
- Times are accurate to the caption line, not to the word.
- It needs a key and a call to our API. If everything must stay on your network, run the library with your own proxies.
Switching from the library #
Before:
from youtube_transcript_api import YouTubeTranscriptApi
snippets = YouTubeTranscriptApi().fetch("jwnez8HdN7E")
text = " ".join(s.text for s in snippets)After:
# pip install cleanscript-ai
from cleanscript_ai import CleanScript
client = CleanScript() # reads CLEANSCRIPT_API_KEY
text = client.transcript("jwnez8HdN7E").textFor timed lines like the library's snippets, add include=["lines"]: each line has start, end and text (an end time instead of a duration). The timestamps guide shows them.
Sources
Spot something wrong? Tell us.