Retrieval works best with chunks that make sense on their own and point back to their source. A transcript request already gives you those. Its paragraphs follow the talk's own breaks instead of a fixed character count, and each has a start time.
So chunking takes a short loop over the paragraphs.
One chunk per paragraph #
This turns a transcript into records for an embedding model or a search index:
# pip install cleanscript-ai
from cleanscript_ai import CleanScript
client = CleanScript()
result = client.transcript("https://www.tiktok.com/@lemondefr/video/7449782721947241750")
post = result.post
def moment(start):
"""YouTube links can open at a time; TikTok and Instagram links can't."""
if post.platform == "youtube":
return f"https://youtu.be/{post.id}?t={int(start)}"
return post.url
chunks = [
{
"text": paragraph.text,
"section": section.title,
"start": paragraph.start,
"url": moment(paragraph.start),
"source": post.url,
"author": post.author.name,
}
for section in result.sections
for paragraph in section.paragraphs
]Le Monde's TikTok on how it makes its explainer videos runs 281 seconds. It comes back as 2 sections and 8 paragraphs of 365 to 919 characters. One record:
{
"text": "Mais alors, on commence par eux. Comment est-ce qu'on choisit nos sujets ? Tous les matins …",
"section": "Dans les coulisses des vidéos",
"start": 29.44,
"url": "https://www.tiktok.com/@lemondefr/video/7449782721947241750",
"source": "https://www.tiktok.com/@lemondefr/video/7449782721947241750",
"author": "Le Monde"
}Embed text and keep the rest as metadata. When your app answers from a chunk, show url so the reader can check it. For a YouTube video, url opens at the paragraph's start.
A post under three minutes usually comes back as one section titled with the post's title. TikTok and Instagram posts have no title, so there section is None.
Bigger or smaller chunks #
Paragraphs run from a few hundred to about a thousand characters, which suits most embedding models. To change that:
- Bigger: join neighbouring paragraphs in the same section up to your limit, and keep the first one's
start. - Smaller: ask for
include=["lines"]and group caption lines instead; each has its own times. See timestamps in JSON.
Add the title and author to the embedded text if questions will name them. A paragraph alone may not say which video it is from.
Many videos #
GET /v1/posts lists a creator's recent posts, newest first; send each post's url to the transcript call. Pass since to get only posts from a date onward, and skip ids you have already indexed. The same request within 10 minutes returns the same list for free. See Creator posts.
When you need facts #
Retrieval finds passages. If you ask the same questions of every video (who sponsors it, what price is quoted), extract fields instead: each value comes with its quote and a link to the moment. See extract fields from a video.
Spot something wrong? Tell us.