Skip to content

Transcripts

Get a post's words as clean text, paragraphs and sections, in the language you read.

View as Markdown

Ask for a transcript

Send the post link as url, with GET and query parameters or POST and a JSON body. Both return the same thing.

curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://www.tiktok.com/@nytimes/video/7489936234010135854" \
  -d language=fr -d include=lines

url takes any link form: watch, Shorts, youtu.be and embed links, TikTok video, photo and share links, Instagram post, Reel and TV links, with or without https://, or a bare YouTube video ID.

What comes back

  • text: the whole transcript, paragraphs separated by blank lines. Often all you need.
  • sections: the creator's chapters, or sections we title from the content (one section for a short post), each with start, end, a url to that moment and its paragraphs.
  • lines: caption-sized lines with their times, with include: ["lines"]: for subtitles, or to find the moment something is said.
  • source and language: where the words came from (below) and their language.
  • warnings: only when the result differs from what you asked: translated, no_speech, uncorrected when some lines couldn't be corrected, or incomplete when part of the post couldn't be read (a slide, the video, or sung lines that may remain as speech). Ask again to get all of it; that request is charged again.
  • post: title, description, author, date, duration and stats, the same object on every route.

Where the words come from

source Meaning
creator The creator's own captions, returned as written.
auto The platform's automatic captions, corrected: misheard words, names and numbers fixed (names spelled as the post's title and description spell them), punctuation added.
transcribed The post had no usable captions, so we transcribed its audio, then corrected it.
translated A translation into the language you asked for: the platform's own when it has one, otherwise ours. A translated warning names the original language.
none No speech (music only, or silent). text is empty, a no_speech warning says so, and on_screen_text holds the text shown in the video.

Posts without captions are transcribed up to 15 minutes long. A longer post without captions returns no_captions.

Languages

Leave language out to get the language spoken in the post. Set it (en, fr, pt-BR) to read in another language: you get the post's own captions in that language when they exist, otherwise a translation (the platform's, or ours), line by line with the same times. A language we can't translate into returns language_unavailable, with the post's languages in details.available.

Times are seconds from the start of the post. Each line starts and ends when its caption does, so times are accurate to the line, not to the word; automatic YouTube captions roll, so a line can start before the previous one ends. A section's url opens YouTube at that moment; TikTok and Instagram links can't open at a moment, so there url is the post link and start gives the moment.

On-screen text

on_screen_text lists the text shown in the video with its times (by slide for photos and carousels); text shown again later is listed again. It is filled automatically when a post has no speech. With speech, ask for it with include: ["on_screen_text"], which adds 1 credit, only when all of it was read.

Cost

1 credit per started 10 minutes of video, or 1 per started minute without captions; our translation adds 1 per started 10 minutes. See Credits & limits.