On this page
YouTube's automatic captions are made for the screen. As text they are hard to read: no capitals, no punctuation, lines that break mid-sentence, and names heard as the nearest common words. YouTube's own help page says they "might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise".
CleanScript corrects the caption text itself. It doesn't retranscribe the audio, and each line keeps the times YouTube gave it.
Before and after #
YouTube's automatic captions for Steve Jobs' 2005 Stanford Commencement Address, from 1:00:
I dropped out of Reed College after the
first 6 months but then stayed around as
a drop in for another 18 months or so
before I really
quit so why' I drop
out it started before I was
born my biological mother was a young
unwed graduate student and she decided
to put me up for
adoption she felt very strongly that IThe same passage after correction:
I dropped out of Reed College after the first 6 months, but then stayed around as a drop-in for another 18 months or so before I really quit. So why'd I drop out? It started before I was born. My biological mother was a young, unwed graduate student, and she decided to put me up for adoption. She felt very strongly that I should be adopted by college graduates.
Punctuation and capitals are back, sentences read across line breaks, and "why' I drop" is "why'd I drop".
This video also has Stanford's own captions, so a normal request returns those, as written. We used its automatic track to measure the correction.
Misheard words #
Some fixes go beyond punctuation. Real ones from the videos we measured:
| Automatic caption | Corrected | Video |
|---|---|---|
| I write the blog weight but why | Wait But Why | Tim Urban, TED |
| gold, frankincense, and mayo | myrrh | Sir Ken Robinson, TED |
| optimistic neoism | nihilism | Kurzgesagt – In a Nutshell |
| Please welcome Ted Harrington | Kit Harington | The Late Show with Stephen Colbert |
What we measured #
We took 15 public videos in English, French, Spanish and German that have both automatic captions and captions their creator wrote. We compared the automatic captions word by word with the creator's, before and after correction, ignoring capitals and punctuation:
- 407 words fixed: wrong in the automatic captions, right after correction.
- 70 words broken: right before, wrong after.
Some errors remain. Correction never changes a number or a negation such as "not". Creator captions are never corrected.
How to tell what you got #
The source field says it:
source |
Meaning |
|---|---|
creator |
The creator's own captions, as written |
auto |
YouTube's automatic captions, corrected |
transcribed |
No usable captions, so we transcribed the audio, then corrected it |
translated |
Our translation, with a warning naming the original language |
# pip install cleanscript-ai
from cleanscript_ai import CleanScript
client = CleanScript()
result = client.transcript("https://youtu.be/jwnez8HdN7E")
print(result.source) # "auto" for this video
print(result.text[:200])If some lines couldn't be corrected, warnings has one with the code uncorrected; ask again for all of it (that request is charged again). No warning means correction covered every line.
When to rely on it #
Corrected text is right for reading, summaries, search and notes. It is not a verbatim record. Before you publish a quote, open the link to its moment and listen.
Next steps #
- Get the transcript: YouTube transcript API.
- Export it: SRT, Markdown or plain text.
- Compare options: youtube-transcript-api alternatives.
Sources
Spot something wrong? Tell us.