# YouTube transcript API: clean text and timestamps

> Get a YouTube video's transcript as JSON in one request, from curl, Python or JavaScript: what comes back, and how languages and missing captions work.

Updated 2026-10-06. HTML: /blog/youtube-transcript-api

A YouTube transcript API takes a video link and returns what is said in it as text. You send one request and get JSON back. The fetching from YouTube happens on the API's side, not from your server.

CleanScript's version is one `GET`. You get the transcript in paragraphs with a link to each section, the video's details, and a `source` field that says where the words came from. TikTok and Instagram links work the same way.

## Make the request

Create a key on the [API keys page](/dashboard/api-keys). New accounts get free credits, no card.

```bash
curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://youtu.be/jwnez8HdN7E"
```

Any link form works: `watch?v=`, `youtu.be`, `/shorts/`, `/live/`, embed links, or a bare video ID.

The SDKs read your key from `CLEANSCRIPT_API_KEY`:

```python
# pip install cleanscript-ai
from cleanscript_ai import CleanScript

client = CleanScript()
result = client.transcript("https://youtu.be/jwnez8HdN7E")
print(result.text)
```

```javascript
// npm install cleanscript
import { CleanScript } from "cleanscript"

const client = new CleanScript()
const result = await client.transcript("https://youtu.be/jwnez8HdN7E")
console.log(result.text)
```

## What comes back

The response for Fireship's [Microsoft's new chip looks like science fiction…](https://www.youtube.com/watch?v=jwnez8HdN7E), shortened:

```json
{
  "post": {
    "platform": "youtube",
    "type": "video",
    "id": "jwnez8HdN7E",
    "url": "https://www.youtube.com/watch?v=jwnez8HdN7E",
    "title": "Microsoft’s new chip looks like science fiction…",
    "author": { "name": "Fireship", "handle": "@Fireship", "url": "https://www.youtube.com/@Fireship" },
    "published_at": "2025-02-21",
    "duration": 259.0,
    "stats": { "views": 2331993, "likes": 71026, "comments": 3300, "shares": null, "saves": null }
  },
  "language": "en",
  "source": "auto",
  "text": "Out of nowhere, Microsoft just announced an impossible new quantum computing chip named Majorana 1, but it's not your average quantum chip. …",
  "sections": [
    {
      "title": "Microsoft's quantum breakthrough",
      "start": 2.19,
      "end": 58.11,
      "url": "https://youtu.be/jwnez8HdN7E?t=2",
      "paragraphs": [{ "start": 2.19, "end": 58.11, "text": "Out of nowhere, Microsoft just announced …" }]
    }
  ],
  "lines": null,
  "on_screen_text": null,
  "warnings": []
}
```

- **`text`**: the whole transcript, paragraphs separated by blank lines. Often all you need.
- **`sections`**: the creator's chapters, or sections we title from the content. Each has a `url` that opens the video at its start, and paragraphs with their own times in seconds.
- **`post`**: the video's details. Every route returns the same object, on every platform. A count the platform doesn't show is `null`, never `0`.
- **`source`**: `creator` (the creator's own captions, as written), `auto` (YouTube's automatic captions, corrected), `transcribed` (no usable captions, so we transcribed the audio), `translated`, or `none` (no speech).
- **`lines`**: caption-sized lines with times, only when you ask with `include=lines`. See [timestamps in JSON](/blog/youtube-transcript-json-timestamps).
- **`warnings`**: empty unless the result differs from what you asked, for example a translation.

Automatic captions are corrected for punctuation, capitals and misheard words. Captions a person wrote come back as written. See [how auto captions get cleaned](/blog/clean-youtube-auto-captions).

## Choose the language

Without `language` you get the language spoken in the video. Add `language=fr` to read in French. If the video has French captions, you get them. If not, you get our translation, with `source` set to `translated` and a warning that names the original language, for 1 credit more per started 10 minutes.

## When a video has no captions

A missing caption track is not an error. We transcribe the audio, correct it, and return it with `source: "transcribed"`. That works for videos up to 15 minutes; a longer one without captions returns `no_captions`. A video with no speech returns empty `text` and a `no_speech` warning.

## What it costs

- New accounts get 100 free credits.
- A transcript costs 1 credit per started 10 minutes of video, so a 25-minute video is 3 credits. Videos can be up to 2 hours.
- A video without captions is transcribed from its audio for 1 credit per started minute instead, so a 5-minute one is 5 credits.
- The same request again within 6 hours is free, and a failed request costs nothing.

Every response carries `X-Credits-Used` and `X-Credits-Remaining`, and `GET /v1/account` shows your balance for free. Packs are on the [pricing page](/pricing).

## Next steps

- Get fields instead of the whole text: [extract fields from a video](/blog/json-schema-fields-from-video).
- TikTok and Instagram links: [TikTok transcript API](/blog/tiktok-transcript-api) and [Instagram Reels transcript](/blog/instagram-reels-transcript).
- Run it on a server: [what to do when a request fails](/blog/youtube-transcripts-in-production).
- Compare with the open-source library: [youtube-transcript-api alternatives](/blog/youtube-transcript-api-alternatives).
- Every field and error: the [API reference](/docs/api-reference).
