# Transcripts

Get a post's words as clean text, paragraphs and sections, in the language you read.

## Ask for a transcript

Send the post link as `url`, with `GET` and query parameters or `POST` and a JSON body. Both return the same thing.

```bash
curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://www.tiktok.com/@nytimes/video/7489936234010135854" \
  -d language=fr -d include=lines
```

```python
result = client.transcript(
    "https://www.tiktok.com/@nytimes/video/7489936234010135854",
    language="fr",
    include=["lines"],
)
```

```typescript
const result = await client.transcript("https://www.tiktok.com/@nytimes/video/7489936234010135854", {
  language: "fr",
  include: ["lines"],
})
```

`url` takes any link form: watch, Shorts, `youtu.be` and embed links, TikTok video, photo and share links, Instagram post, Reel and TV links, with or without `https://`, or a bare YouTube video ID.

## What comes back

- `text`: the whole transcript, paragraphs separated by blank lines. Often all you need.
- `sections`: the creator's chapters, or sections we title from the content (one section for a short post), each with `start`, `end`, a `url` to that moment and its `paragraphs`.
- `lines`: caption-sized lines with their times, with `include: ["lines"]`: for subtitles, or to find the moment something is said.
- `source` and `language`: where the words came from (below) and their language.
- `warnings`: only when the result differs from what you asked: `translated`, `no_speech`, `uncorrected` when some lines couldn't be corrected, or `incomplete` when part of the post couldn't be read (a slide, the video, or sung lines that may remain as speech). Ask again to get all of it; that request is charged again.
- `post`: title, description, author, date, duration and stats, the same object on every route.

## Where the words come from

| `source` | Meaning |
| --- | --- |
| `creator` | The creator's own captions, returned as written. |
| `auto` | The platform's automatic captions, corrected: misheard words, names and numbers fixed (names spelled as the post's title and description spell them), punctuation added. |
| `transcribed` | The post had no usable captions, so we transcribed its audio, then corrected it. |
| `translated` | A translation into the `language` you asked for: the platform's own when it has one, otherwise ours. A `translated` warning names the original language. |
| `none` | No speech (music only, or silent). `text` is empty, a `no_speech` warning says so, and `on_screen_text` holds the text shown in the video. |

Posts without captions are transcribed up to 15 minutes long. A longer post without captions returns `no_captions`.

## Languages

Leave `language` out to get the language spoken in the post. Set it (`en`, `fr`, `pt-BR`) to read in another language: you get the post's own captions in that language when they exist, otherwise a translation (the platform's, or ours), line by line with the same times. A language we can't translate into returns `language_unavailable`, with the post's languages in `details.available`.

## Times and links

Times are seconds from the start of the post. Each line starts and ends when its caption does, so times are accurate to the line, not to the word; automatic YouTube captions roll, so a line can start before the previous one ends. A section's `url` opens YouTube at that moment; TikTok and Instagram links can't open at a moment, so there `url` is the post link and `start` gives the moment.

## On-screen text

`on_screen_text` lists the text shown in the video with its times (by `slide` for photos and carousels); text shown again later is listed again. It is filled automatically when a post has no speech. With speech, ask for it with `include: ["on_screen_text"]`, which adds 1 credit, only when all of it was read.

## Cost

1 credit per started 10 minutes of video, or 1 per started minute without captions; our translation adds 1 per started 10 minutes. See [Credits & limits](https://cleanscript.ai/docs/credits.md).
