To get a YouTube transcript as JSON with timestamps, ask for lines. Each caption line comes back with its start and end in seconds, ready to index, search or turn into subtitles.
curl -G https://api.cleanscript.ai/v1/transcript \
-H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
--data-urlencode "url=https://youtu.be/jwnez8HdN7E" \
-d include=lines# pip install cleanscript-ai
from cleanscript_ai import CleanScript
client = CleanScript()
result = client.transcript("https://youtu.be/jwnez8HdN7E", include=["lines"])
for line in result.lines:
print(f"{line.start:7.2f} {line.text}")// npm install cleanscript
import { CleanScript } from "cleanscript"
const client = new CleanScript()
const result = await client.transcript("https://youtu.be/jwnez8HdN7E", { include: ["lines"] })
for (const line of result.lines ?? []) console.log(line.start.toFixed(2), line.text)What the lines look like #
TikTok and Instagram links work the same way. The first lines of Le Monde's TikTok about its explainer videos, from TikTok's automatic captions:
"lines": [
{ "start": 0.08, "end": 4.56, "text": "Comment réalisons nos vidéos verticales au Monde ? Bienvenue dans notre open space." },
{ "start": 4.68, "end": 8.88, "text": "C'est ici qu'on fabrique les vidéos d'explication qui sont publiées sur les réseaux sociaux du journal" },
{ "start": 9.04, "end": 12.76, "text": "le monde. On existe depuis deux-mille-seize, et avant ça ressemblait à ça," }
]Times are seconds from the start, to two decimals. A line starts and ends when its caption does, so times are accurate to the line, not the word: good for finding a moment, not for cutting on a syllable. YouTube's automatic captions roll, so a line can start before the previous one ends.
Automatic captions are corrected, which is why these lines have capitals and punctuation. Captions a person wrote come back as written. source tells you which you got.
Lines or paragraphs #
Without include, you get text and sections, whose paragraphs have their own start and end times. Paragraphs are better for reading and for a model. Lines are better for subtitles, search and pointing at a short span. With include=lines you get both.
Link to the exact moment #
Each section has a url that opens the video at its start. For a single line, build the link yourself, rounding down to a whole second so the player lands just before the words:
post = result.post
for line in result.lines[:3]:
print(f"https://youtu.be/{post.id}?t={int(line.start)}", line.text)TikTok and Instagram links can't open at a moment, so there, link the post and give the time as text.
Next steps #
- Turn the lines into SRT, Markdown or plain text.
- Split a transcript for search: chunks for RAG.
- Every field: Transcripts and the API reference.
Spot something wrong? Tell us.