# Chunk a video transcript for RAG, with a link to each moment

> Turn a YouTube, TikTok or Instagram video into search-ready RAG chunks that keep author, section and a link to the moment, in a short Python loop.

Updated 2026-10-06. HTML: /blog/transcript-chunks-for-rag

Retrieval works best with chunks that make sense on their own and point back to their source. A transcript request already gives you those. Its paragraphs follow the talk's own breaks instead of a fixed character count, and each has a start time.

So chunking takes a short loop over the paragraphs.

## One chunk per paragraph

This turns a transcript into records for an embedding model or a search index:

```python
# pip install cleanscript-ai
from cleanscript_ai import CleanScript

client = CleanScript()
result = client.transcript("https://www.tiktok.com/@lemondefr/video/7449782721947241750")
post = result.post

def moment(start):
    """YouTube links can open at a time; TikTok and Instagram links can't."""
    if post.platform == "youtube":
        return f"https://youtu.be/{post.id}?t={int(start)}"
    return post.url

chunks = [
    {
        "text": paragraph.text,
        "section": section.title,
        "start": paragraph.start,
        "url": moment(paragraph.start),
        "source": post.url,
        "author": post.author.name,
    }
    for section in result.sections
    for paragraph in section.paragraphs
]
```

[Le Monde's TikTok](https://www.tiktok.com/@lemondefr/video/7449782721947241750) on how it makes its explainer videos runs 281 seconds. It comes back as 2 sections and 8 paragraphs of 365 to 919 characters. One record:

```json
{
  "text": "Mais alors, on commence par eux. Comment est-ce qu'on choisit nos sujets ? Tous les matins …",
  "section": "Dans les coulisses des vidéos",
  "start": 29.44,
  "url": "https://www.tiktok.com/@lemondefr/video/7449782721947241750",
  "source": "https://www.tiktok.com/@lemondefr/video/7449782721947241750",
  "author": "Le Monde"
}
```

Embed `text` and keep the rest as metadata. When your app answers from a chunk, show `url` so the reader can check it. For a YouTube video, `url` opens at the paragraph's start.

A post under three minutes usually comes back as one section titled with the post's title. TikTok and Instagram posts have no title, so there `section` is `None`.

## Bigger or smaller chunks

Paragraphs run from a few hundred to about a thousand characters, which suits most embedding models. To change that:

- **Bigger:** join neighbouring paragraphs in the same section up to your limit, and keep the first one's `start`.
- **Smaller:** ask for `include=["lines"]` and group caption lines instead; each has its own times. See [timestamps in JSON](/blog/youtube-transcript-json-timestamps).

Add the title and author to the embedded text if questions will name them. A paragraph alone may not say which video it is from.

## Many videos

`GET /v1/posts` lists a creator's recent posts, newest first; send each post's `url` to the transcript call. Pass `since` to get only posts from a date onward, and skip `id`s you have already indexed. The same request within 10 minutes returns the same list for free. See [Creator posts](/docs/posts).

## When you need facts

Retrieval finds passages. If you ask the same questions of every video (who sponsors it, what price is quoted), extract fields instead: each value comes with its quote and a link to the moment. See [extract fields from a video](/blog/json-schema-fields-from-video).
