# CleanScript docs Every page of https://cleanscript.ai/docs in one file. Index: https://cleanscript.ai/llms.txt. OpenAPI: https://cleanscript.ai/openapi.json. --- # CleanScript docs Source: https://cleanscript.ai/docs CleanScript is one HTTP API. Send a post link; get back its words as clean text in paragraphs and sections, or JSON fields with a quote for every value. ```bash curl -G https://api.cleanscript.ai/v1/transcript \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "url=https://youtu.be/jwnez8HdN7E" ``` ## What you can get - [**A transcript**](https://cleanscript.ai/docs/transcripts.md): `GET /v1/transcript` returns clean text in paragraphs and sections, with times and a link to each moment. Automatic captions are corrected; posts without captions are transcribed. - [**Fields**](https://cleanscript.ai/docs/extract.md): `POST /v1/extract` takes a question, a preset or your own JSON Schema, and returns JSON with a quote and the moment for every value. - [**A creator's posts**](https://cleanscript.ai/docs/posts.md): `GET /v1/posts` lists recent posts, ready for the two calls above. - **Your balance**: `GET /v1/account`, free. ## The basics - **Base URL:** `https://api.cleanscript.ai` - **Auth:** `Authorization: Bearer `. Create a key in the [dashboard](https://cleanscript.ai/dashboard/api-keys); new accounts get 100 free credits, no card. - **Posts:** YouTube videos and Shorts, TikTok posts, Instagram posts and Reels (TikTok and Instagram are in beta). Any link form works, or a bare YouTube video ID. - **Shape:** JSON in and out, snake_case, times in seconds. Every field is always present; unknown values are `null`. - **Errors:** one shape, never charged. See [errors](https://cleanscript.ai/docs/credits.md#errors). ## Ways in - **Code:** curl, or the Python and TypeScript SDKs. Start with the [Quickstart](https://cleanscript.ai/docs/quickstart.md). - **AI agents:** the MCP server at `https://api.cleanscript.ai/mcp`, a skill for coding agents, and `llms.txt`. See [Agents & MCP](https://cleanscript.ai/docs/agents.md). - **No-code tools:** n8n, Make and Zapier over HTTP. See [Batches & automations](https://cleanscript.ai/docs/automations.md). - **Your browser:** the [Playground](https://cleanscript.ai/playground) runs a post and shows the request behind it. --- # Quickstart Source: https://cleanscript.ai/docs/quickstart **1. Create a key.** Sign in, open [API keys](https://cleanscript.ai/dashboard/api-keys) and create one. New accounts get 100 free credits, no card. **2. Send a request** with the key in `CLEANSCRIPT_API_KEY`: ```bash export CLEANSCRIPT_API_KEY="" curl -G https://api.cleanscript.ai/v1/transcript \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "url=https://youtu.be/jwnez8HdN7E" ``` ```python # pip install cleanscript-ai from cleanscript_ai import CleanScript client = CleanScript() # reads CLEANSCRIPT_API_KEY print(client.transcript("https://youtu.be/jwnez8HdN7E").text) ``` ```typescript // npm install cleanscript import { CleanScript } from "cleanscript" const client = new CleanScript() // reads CLEANSCRIPT_API_KEY console.log((await client.transcript("https://youtu.be/jwnez8HdN7E")).text) ``` **3. Read the result.** `text` is the whole transcript; `sections` holds it as paragraphs with their times and a link to each moment. This 4-minute video costs 1 credit, and the same request again within 6 hours is free. ```json Response (shortened) {"post":{"platform":"youtube","duration":259},"language":"en","source":"auto","text":"Out of nowhere, Microsoft just announced…","sections":[{"title":"Microsoft's quantum breakthrough","start":2.19,"end":58.11,"url":"https://youtu.be/jwnez8HdN7E?t=2","paragraphs":[{"start":2.19,"end":58.11,"text":"…"}]}],"warnings":[]} ``` ## Next - Ask for fields instead of text: `POST /v1/extract` with `{"url": "…", "prompt": "What does it say about pricing?"}`. See [Extract fields](https://cleanscript.ai/docs/extract.md). - Use it from Claude, ChatGPT, Cursor or a coding agent: [Agents & MCP](https://cleanscript.ai/docs/agents.md). - Every parameter and field: [API reference](https://cleanscript.ai/docs/api-reference.md). --- # Agents & MCP Source: https://cleanscript.ai/docs/agents ## The MCP server Add `https://api.cleanscript.ai/mcp` to any client that supports remote MCP servers (streamable HTTP). You sign in to CleanScript once and approve the connection; clients that can't sign in send an API key instead, as `Authorization: Bearer `. The [Agents page](https://cleanscript.ai/connect) sets up each client in one step. | Tool | What it does | Credits | | --- | --- | --- | | `get_transcript` | The transcript as paragraphs under their section titles, each with its `[m:ss]` time. Long videos come in pages: pass `next_start` as `start`. | Same as `/v1/transcript` | | `extract` | Fields from a `prompt`, a `preset` or a `schema`, with quotes. | Same as `/v1/extract` | | `list_posts` | A creator's recent posts (`author`, `platform`, `limit`, `since`). | 1 credit | | `check_account` | Checks the connection and the credits left. | Free | Results come back in the tool reply, sized for a model's context. Calling `get_transcript` or `extract` again for the same thing within 6 hours is free (`list_posts`: within 10 minutes), and a call that fails is never charged. ## Set up your client **Claude** (web and desktop): open Settings, then Connectors, add a custom connector with the URL above, and sign in. **ChatGPT**: turn on developer mode, add a connector with the URL above, and sign in. **Claude Code**: ```bash claude mcp add --transport http cleanscript https://api.cleanscript.ai/mcp ``` Then run `/mcp` in Claude Code and sign in. To use an API key instead, add `--header "Authorization: Bearer $CLEANSCRIPT_API_KEY"`. **Codex**: ```bash codex mcp add cleanscript --url https://api.cleanscript.ai/mcp codex mcp login cleanscript ``` **Cursor**, in `.cursor/mcp.json` (or `~/.cursor/mcp.json` for every project): ```json .cursor/mcp.json {"mcpServers":{"cleanscript":{"url":"https://api.cleanscript.ai/mcp"}}} ``` **VS Code**, in `.vscode/mcp.json`: ```json .vscode/mcp.json {"servers":{"cleanscript":{"type":"http","url":"https://api.cleanscript.ai/mcp"}}} ``` ## Build it into a project To have a coding agent add CleanScript to your code, give it the skill. For Claude Code: ```bash mkdir -p .claude/skills/cleanscript curl -so .claude/skills/cleanscript/SKILL.md https://cleanscript.ai/SKILL.md ``` Or paste this prompt into any coding agent: ```text Prompt Add CleanScript to this project so I can get transcripts and fields from YouTube, TikTok and Instagram links. Read https://cleanscript.ai/SKILL.md first. The API is https://api.cleanscript.ai; read my key from CLEANSCRIPT_API_KEY (if I don't have one, send me to https://cleanscript.ai/dashboard/api-keys). Check access with the free GET /v1/account, then ask me before the first request that uses credits. Tell me what you built and how to run it. ``` ## Docs for agents - [`/llms.txt`](https://cleanscript.ai/llms.txt): the base URL, auth, two calls and an index of these pages. - [`/llms-full.txt`](https://cleanscript.ai/llms-full.txt): every page in one file, about 40 KB. - Every page as Markdown: add `.md` to its address ([`/docs/quickstart.md`](https://cleanscript.ai/docs/quickstart.md)), or request it with `Accept: text/markdown`. - [`/SKILL.md`](https://cleanscript.ai/SKILL.md): the skill for coding agents. - [`/openapi.json`](https://cleanscript.ai/openapi.json): the OpenAPI 3.1 description, with examples. - [`/integration.json`](https://cleanscript.ai/integration.json): routes, presets, credits and links as one JSON object. --- # Transcripts Source: https://cleanscript.ai/docs/transcripts ## Ask for a transcript Send the post link as `url`, with `GET` and query parameters or `POST` and a JSON body. Both return the same thing. ```bash curl -G https://api.cleanscript.ai/v1/transcript \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "url=https://www.tiktok.com/@nytimes/video/7489936234010135854" \ -d language=fr -d include=lines ``` ```python result = client.transcript( "https://www.tiktok.com/@nytimes/video/7489936234010135854", language="fr", include=["lines"], ) ``` ```typescript const result = await client.transcript("https://www.tiktok.com/@nytimes/video/7489936234010135854", { language: "fr", include: ["lines"], }) ``` `url` takes any link form: watch, Shorts, `youtu.be` and embed links, TikTok video, photo and share links, Instagram post, Reel and TV links, with or without `https://`, or a bare YouTube video ID. ## What comes back - `text`: the whole transcript, paragraphs separated by blank lines. Often all you need. - `sections`: the creator's chapters, or sections we title from the content (one section for a short post), each with `start`, `end`, a `url` to that moment and its `paragraphs`. - `lines`: caption-sized lines with their times, with `include: ["lines"]`: for subtitles, or to find the moment something is said. - `source` and `language`: where the words came from (below) and their language. - `warnings`: only when the result differs from what you asked: `translated`, `no_speech`, `uncorrected` when some lines couldn't be corrected, or `incomplete` when part of the post couldn't be read (a slide, the video, or sung lines that may remain as speech). Ask again to get all of it; that request is charged again. - `post`: title, description, author, date, duration and stats, the same object on every route. ## Where the words come from | `source` | Meaning | | --- | --- | | `creator` | The creator's own captions, returned as written. | | `auto` | The platform's automatic captions, corrected: misheard words, names and numbers fixed (names spelled as the post's title and description spell them), punctuation added. | | `transcribed` | The post had no usable captions, so we transcribed its audio, then corrected it. | | `translated` | A translation into the `language` you asked for: the platform's own when it has one, otherwise ours. A `translated` warning names the original language. | | `none` | No speech (music only, or silent). `text` is empty, a `no_speech` warning says so, and `on_screen_text` holds the text shown in the video. | Posts without captions are transcribed up to 15 minutes long. A longer post without captions returns `no_captions`. ## Languages Leave `language` out to get the language spoken in the post. Set it (`en`, `fr`, `pt-BR`) to read in another language: you get the post's own captions in that language when they exist, otherwise a translation (the platform's, or ours), line by line with the same times. A language we can't translate into returns `language_unavailable`, with the post's languages in `details.available`. ## Times and links Times are seconds from the start of the post. Each line starts and ends when its caption does, so times are accurate to the line, not to the word; automatic YouTube captions roll, so a line can start before the previous one ends. A section's `url` opens YouTube at that moment; TikTok and Instagram links can't open at a moment, so there `url` is the post link and `start` gives the moment. ## On-screen text `on_screen_text` lists the text shown in the video with its times (by `slide` for photos and carousels); text shown again later is listed again. It is filled automatically when a post has no speech. With speech, ask for it with `include: ["on_screen_text"]`, which adds 1 credit, only when all of it was read. ## Cost 1 credit per started 10 minutes of video, or 1 per started minute without captions; our translation adds 1 per started 10 minutes. See [Credits & limits](https://cleanscript.ai/docs/credits.md). --- # Extract fields Source: https://cleanscript.ai/docs/extract ## Three ways to ask `POST /v1/extract` takes the post `url` and one of these: - `prompt`: a question or instruction in plain words, such as "What does it say about pricing?". - `preset`: one of four ready-made schemas, `summary`, `mentions`, `ad-breakdown`, `claims`. - `schema`: your own JSON Schema. A `prompt` can also go with a `preset` or `schema`, as extra guidance: "focus on the sponsor segment". ```bash curl https://api.cleanscript.ai/v1/extract \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ -H "Content-Type: application/json" \ -d '{"url": "https://youtu.be/jwnez8HdN7E", "prompt": "What does it say about how the chip works?"}' ``` ```python result = client.extract("https://youtu.be/jwnez8HdN7E", prompt="What does it say about how the chip works?") print(result.data["answer"]) ``` ```typescript const result = await client.extract("https://youtu.be/jwnez8HdN7E", { prompt: "What does it say about how the chip works?" }) console.log(result.data.answer) ``` A `prompt` on its own returns `data` with an `answer` (`null` when the post doesn't address it) and the `points` it rests on (shortened): ```json Response {"data":{"answer":"The chip uses Majorana zero modes at the ends of atomically engineered nanowires, which Microsoft says are resistant to quantum decoherence. It computes by braiding and fusing these particles, then measuring whether the wire contains an even or odd number of electrons; the wires are made into a “topoconductor” and can be linked together. The chip must be kept near absolute zero, and its claimed ability to scale to millions of qubits has not yet been demonstrated.","points":["The approach is based on Majorana fermions, described as particles that are their own antiparticles and highly resistant to decoherence."]},"citations":{"points[0]":[{"quote":"It's based on the Majorana fermion, which is a subatomic particle that's also its own antiparticle.","source":"speech","start":96.44,"end":102.06,"url":"https://youtu.be/jwnez8HdN7E?t=96"}]},"removed":[],"warnings":[]} ``` ## What comes back - `data`: your fields, matching the schema. Any value can be `null` when the post doesn't support it, so a schema never fails the request. - `citations`: the quotes behind each value, keyed by its path. - `removed`: paths of values we dropped because no quote in the post supports them. - `warnings`: only when the result differs from the request: `incomplete` when part of the post couldn't be read, or `uncorrected`. - `post`: the post, as on every route. Add `include: ["transcript"]` to get the transcript the values come from. ## Quotes Every value comes with up to 3 quotes. A quote has the words as they appear in the post (`quote`; a spoken quote is widened to the whole sentences it is part of), where they appear (`source`: `speech`, `title`, `description` or `on_screen`), and, for speech and on-screen text in a video, its `start`, `end` and a `url` to that moment. - Keys are readable paths: `offer`, `claims[1]`, `host.name`. An object in a list is quoted once as a whole: `sponsors[0]`. - Neighbouring caption lines are joined into one quote. - `null`, `[]`, `false` and `0` need no quote. A `true` always has one; a yes/no the post doesn't support comes back `false`. - Values are checked against the post. A value no quote supports is removed and listed in `removed`. - Values are written in the `language` you ask for (by default, the transcript's); quotes stay word for word. ## Presets Each preset is a JSON Schema published on the site. Sending `"preset": "summary"` is the same as sending that file as `schema`; its field descriptions carry all the guidance. | Preset | For | Fields | | --- | --- | --- | | [`summary`](https://cleanscript.ai/schemas/summary.json) | What a post says, for notes, newsletters and research | `summary`, `key_points`, `quotes`, `people` | | [`mentions`](https://cleanscript.ai/schemas/mentions.json) | Brands, products and sponsors in a post, for brand monitoring and influencer marketing | `sponsors`, `mentions`, `recommendations` | | [`ad-breakdown`](https://cleanscript.ai/schemas/ad-breakdown.json) | A strategist's breakdown of an ad or a promotional post, from what is said and the post's text | `hook`, `hook_type`, `angle`, `audience`, `pain_points`, `promise`, `product`, `offer`, `claims`, `proof`, `objections_handled`, `cta`, `structure`, `tone` | | [`claims`](https://cleanscript.ai/schemas/claims.json) | What a post asserts, for research, fact-checking and brand safety | `claims`, `figures`, `sources` | Presets only gain fields; a change that would break one gets a new name. ## Your own schema Send any JSON Schema (draft 2020-12) with an object at the root, up to 32 KB; local `$ref`s work; recursive schemas and `pattern` don't. Describe each field in its `description`: that is what guides the answer. ```json Request {"url":"https://www.youtube.com/watch?v=x7X9w_GIm1s","language":"fr","schema":{"properties":{"creator":{"description":"Who created the language.","type":"string"},"release_year":{"type":"integer"},"use_cases":{"items":{"type":"string"},"type":"array"}},"required":["creator","release_year","use_cases"],"type":"object"}} ``` ## Write fields that work - **Describe each field in plain words**, and say what it is not when two fields could overlap: "the deal offered; not the product". - **Ask for the thing, not its time.** Every value's quote already carries `start`, `end` and a link. - **Keep numbers as stated.** "three to six months" doesn't fit an integer. Use a string for the figure as said and another for what it measures. - **Use an `enum` for anything you will group by** (type, sentiment, stage), with an `"other"` option. - **Need exact words? Use the quote.** A description saying "word for word" is a request; the quote is the guarantee. - **Prefer specific, quotable items** over whole-video judgements, and on long videos ask for "the 10 most important" rather than "every". - **Ask about what is said or written.** Fields are filled from the speech, the title and the description (and the on-screen text when a post has no speech), not from images, music or sounds. - **Give context in the root `description`**, such as "we track mentions of our brand and its competitors". ## Cost 3 credits per started 10 minutes of video, the transcript included. Without captions: the transcription, 1 credit per started minute, plus 2 per started 10 minutes. See [Credits & limits](https://cleanscript.ai/docs/credits.md). --- # Creator posts Source: https://cleanscript.ai/docs/posts ## List a creator's posts `GET /v1/posts` returns a creator's recent posts, newest first. `author` is a profile or channel link in any form (`youtube.com/@name`, `/channel/UC…`, `tiktok.com/@name`, `instagram.com/name`), or a handle such as `@Fireship` with `platform`. ```bash curl -G https://api.cleanscript.ai/v1/posts \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "author=https://www.youtube.com/@Fireship" \ -d limit=5 ``` ```python listing = client.posts("https://www.youtube.com/@Fireship", limit=5) for post in listing.posts: print(post.published_at, post.title, post.url) ``` ```typescript const listing = await client.posts("https://www.youtube.com/@Fireship", { limit: 5 }) for (const post of listing.posts) console.log(post.published_at, post.title, post.url) ``` The response has the creator (`author`, with `followers`) and `posts`. Each post is the same object every route returns, so its `url` goes straight into `/v1/transcript` or `/v1/extract`. ## Options - `limit`: how many posts, 12 by default. One call returns at most 30 on YouTube, 35 on TikTok and 12 on Instagram. - `since`: only posts published on or after a date, `YYYY-MM-DD`. - `platform`: needed only when `author` is a handle. A call costs 1 credit. The list can be up to 10 minutes old. The same call within 10 minutes is free. On YouTube it holds videos, not Shorts, without descriptions, likes or comments, and with rounded view counts (the transcript call returns exact ones), and the dates of older videos are approximate. A profile that doesn't exist or is private returns `author_not_found`, with `details.reason`. ## Check for new posts To follow creators, run a job on a schedule (daily is plenty for most): 1. Call `/v1/posts` with `since` set to the date of your last check. On YouTube, go one day further back, since its dates are approximate. 2. Skip posts whose `id` you have already seen: `since` is a date, so a post from the day of your last check comes back again. 3. Send each new post's `url` to `/v1/transcript` or `/v1/extract`, and store the result with the post `id`. ```python from cleanscript_ai import CleanScript client = CleanScript() seen = set(load_seen_ids()) # your storage for creator in ["https://www.youtube.com/@Fireship", "https://www.tiktok.com/@nytimes"]: listing = client.posts(creator, since=last_check_date) for post in listing.posts: if post.id in seen: continue fields = client.extract(post.url, preset="mentions") save(post, fields) # your storage seen.add(post.id) ``` ```typescript import { CleanScript } from "cleanscript" const client = new CleanScript() const seen = new Set(await loadSeenIds()) // your storage for (const creator of ["https://www.youtube.com/@Fireship", "https://www.tiktok.com/@nytimes"]) { const listing = await client.posts(creator, { since: lastCheckDate }) for (const post of listing.posts) { if (seen.has(post.id)) continue const fields = await client.extract(post.url, { preset: "mentions" }) await save(post, fields) // your storage seen.add(post.id) } } ``` Each check costs 1 credit per creator, plus the transcripts or fields of the new posts. For many creators, see [Batches & automations](https://cleanscript.ai/docs/automations.md). --- # Batches & automations Source: https://cleanscript.ai/docs/automations ## n8n, Make and Zapier Each call is one HTTP step. Set it up like this in any tool: - **Method and URL:** `POST https://api.cleanscript.ai/v1/transcript`, or `/v1/extract` for fields. - **Header:** `Authorization` with the value `Bearer `. Store it as a credential, not in the step. - **Body (JSON):** `{"url": ""}`, plus `prompt`, `preset` or `schema` for fields. - **Timeout:** 2 minutes or more. Short posts answer in seconds; long videos take longer. - **Next steps:** map `text` (transcript) or `data` (fields) from the response. In **n8n**, use the HTTP Request node with a Header Auth credential and Send Body as JSON; set the timeout under Options. In **Make**, use HTTP, Make a request, with Parse response on. In **Zapier**, use Webhooks by Zapier, Custom Request. If a step gives up before the answer arrives, run it again: asking again for the same thing within 6 hours is free, and a request that fails is never charged. ## Many posts from code Your account allows a number of requests a minute (`rate_limit_per_minute` in `GET /v1/account`). Run a few requests at a time; the SDKs wait and retry when you hit the limit (`429`) or a temporary error (`503`), and send an `Idempotency-Key` so a retry is never charged twice. ```python from concurrent.futures import ThreadPoolExecutor from cleanscript_ai import CleanScript, CleanScriptError client = CleanScript(max_retries=4) def run(url): try: return url, client.transcript(url).text except CleanScriptError as error: return url, error # error.code says why: video_unavailable, no_captions… with ThreadPoolExecutor(max_workers=4) as pool: for url, result in pool.map(run, urls): print(url, result) ``` ```typescript import { CleanScript, CleanScriptError } from "cleanscript" const client = new CleanScript({ maxRetries: 4 }) const queue = [...urls] const results: { url: string; text?: string; error?: CleanScriptError }[] = [] await Promise.all( Array.from({ length: 4 }, async () => { for (let url = queue.shift(); url; url = queue.shift()) { try { results.push({ url, text: (await client.transcript(url)).text }) } catch (error) { if (!(error instanceof CleanScriptError)) throw error results.push({ url, error }) // error.code says why } } }) ) ``` ## Handle failures by code - **Retry later:** `rate_limited`, `source_unavailable`, `processing_failed` (the SDKs retry these for you; over HTTP, wait for `Retry-After`) and `timeout` (retry once; the result is often ready by then). - **Skip and record:** `video_unavailable`, `no_captions`, `video_too_long`, `language_unavailable`, `invalid_url`, `unsupported_platform`. Retrying won't change the answer. - **Stop the batch:** `insufficient_credits` (top up, then resume) and `unauthorized` (check the key). Over HTTP, send an `Idempotency-Key` header unique to the post, such as `transcript-` plus the post ID (1 to 128 letters, digits, `.`, `_`, `:` or `-`), and reuse it when you retry that post: within 6 hours a retry returns the first successful result instead of running again. Retrying a request that failed, with the same key, runs it again. ## Before a large batch Check your balance with `GET /v1/account` (free). Every response also carries `X-Credits-Used` and `X-Credits-Remaining`. A 10-minute video costs 1 credit for its transcript, or 3 credits with fields; see [Credits & limits](https://cleanscript.ai/docs/credits.md) for the rates. --- # Credits & limits Source: https://cleanscript.ai/docs/credits ## Credits New accounts get 100 free credits, no card. After that, buy a pack in the [dashboard](https://cleanscript.ai/dashboard/billing); credits are valid for 12 months and add to your balance. | Pack | Credits | Price | Per credit | | --- | --- | --- | --- | | Starter | 1,000 | $10 | 1¢ | | Builder | 5,000 | $40 | 0.8¢ | | Pro | 25,000 | $150 | 0.6¢ | ## What a request costs - **Transcript:** 1 credit per started 10 minutes of video: an 11-minute video is 2 credits. - **Fields** (`/v1/extract`): 3 credits per started 10 minutes, the transcript included. - **Posts without captions:** transcribed for 1 credit per started minute instead: a 45-second Reel is 1 credit. - **A translation the platform doesn't provide** (into `language`): 1 credit more per started 10 minutes. - **On-screen text** alongside speech: 1 credit more. For a post without speech it is the transcript. - **Creator posts** (`/v1/posts`): 1 credit per call. - **Account** (`/v1/account`): free. | Post | Transcript | Transcript and fields | | --- | --- | --- | | A 45-second TikTok or Reel, with or without captions | 1 credit | 3 credits | | A 10-minute YouTube video | 1 credit | 3 credits | | A 15-minute video without captions | 15 credits | 19 credits | | A 1-hour podcast | 6 credits | 18 credits | | A creator's post list (`/v1/posts`) | 1 credit | | Free, always: any request that fails. Asking again for the same thing within 6 hours is free. If you got this post's transcript in that time, asking for its fields costs only the fields. Asking again after a result with an `incomplete` or `uncorrected` warning runs the request again in full and is charged again. Every response says what it used in `X-Credits-Used` and what is left in `X-Credits-Remaining`. With no credits left, every request except `/v1/account` is `insufficient_credits`, even a free repeat. ## Limits - **Requests a minute:** set per account, 10 for new accounts; yours is `rate_limit_per_minute` in `GET /v1/account`. Over it you get `rate_limited` with `Retry-After`. Need more? Write to [mohamed@messaad.dev](mailto:mohamed@messaad.dev). - **Unavailable posts and profiles:** 30 an hour. Requests that end in `video_unavailable`, `author_not_found` or `source_unavailable` are free, but past 30 in an hour you get `rate_limited` until the hour is up. Asking again within 30 minutes for a post or profile that doesn't exist returns the same 404 and doesn't count. - **Post length:** up to 2 hours. Posts without captions are transcribed up to 15 minutes. - **Time:** a request answers within 120 seconds, or returns `timeout`. - **Size:** request body up to 128 KB, schema up to 32 KB, response up to 8 MB. - **Creator posts:** up to 30 per call on YouTube (videos, not Shorts), 35 on TikTok, 12 on Instagram. ## Errors Every error has the same shape. `code` is stable; `message` says what happened in words you can show; `param` names the field at fault; `details` carries data you can act on. Quote `request_id` when you contact us. ```json Error 402 {"error":{"code":"insufficient_credits","message":"This request needs 1 credit and the account has 0 left. Buy more at https://cleanscript.ai/pricing.","details":{"credits_remaining":0,"credits_required":1,"top_up_url":"https://cleanscript.ai/pricing"}},"request_id":"5f0c7d1e-0b7a-4a39-9c1e-3c1f2b8f6a10"} ``` | Code | Status | What to do | | --- | --- | --- | | `invalid_request` | 400 | Fix the field named in `param`; the message says how. Also 404 or 405 for an unknown route or method, and 413 for a body over 128 KB or a result over 8 MB. | | `invalid_url` | 400 | Send a link to one post: a YouTube watch, Shorts or `youtu.be` link, a TikTok video link, an Instagram `/p/` or `/reel/` link. | | `unsupported_platform` | 400 | Only YouTube, TikTok and Instagram posts work. | | `invalid_schema` | 400 | Fix the JSON Schema; the message names the problem. | | `unauthorized` | 401 | Send `Authorization: Bearer ` with a key from the [dashboard](https://cleanscript.ai/dashboard/api-keys) that isn't revoked. | | `insufficient_credits` | 402 | Top up at `details.top_up_url`. `details` also has `credits_required` and `credits_remaining`. | | `video_unavailable` | 404 | `details.reason` is `private`, `removed`, `age_restricted` or `live` on YouTube, and `unavailable` on TikTok and Instagram. A live stream works once it has ended. | | `author_not_found` | 404 | Check the profile link or handle; `details.reason` is `not_found` or `private`. | | `idempotency_conflict` | 409 | The `Idempotency-Key` was used for a different request (use a new key), or the first request is still running (retry after `Retry-After`). | | `no_captions` | 422 | The post has no captions and can't be transcribed: it is over 15 minutes, or its audio isn't available. | | `language_unavailable` | 422 | Choose a language from `details.available`, or leave `language` out. | | `video_too_long` | 422 | The post is over 2 hours. | | `rate_limited` | 429 | Wait for `Retry-After` seconds, then retry. | | `source_unavailable` | 503 | The platform didn't answer. Retry after `Retry-After`. | | `processing_failed` | 503 | Something failed on our side. Retry after `Retry-After`. | | `timeout` | 504 | Retry; the result is often ready by then. | --- # API reference Source: https://cleanscript.ai/docs/api-reference ## Basics - **Base URL:** `https://api.cleanscript.ai`. JSON in and out, snake_case, times in seconds. - **Auth:** `Authorization: Bearer ` on every request. Keys are in the [dashboard](https://cleanscript.ai/dashboard/api-keys). - **Headers on every response:** `X-Request-Id`, `X-Credits-Used`, and `X-Credits-Remaining` once the key is accepted; `Retry-After` on `429` and `503`. - **Retries:** send an optional `Idempotency-Key` (1 to 128 letters, digits, `.`, `_`, `:` or `-`) on `/v1/transcript` and `/v1/extract`. Retrying with the same key and request within 6 hours returns the first successful result, free. After an error, retrying with the same key runs the request again. - **Errors:** `{"error": {"code", "message", "param", "details"}, "request_id"}`. Codes and what to do: [Credits & limits](https://cleanscript.ai/docs/credits.md#errors). - **OpenAPI:** [`/openapi.json`](https://cleanscript.ai/openapi.json). ## GET /v1/transcript A post's words as clean text, paragraphs and sections. `POST /v1/transcript` takes the same parameters as a JSON body. Costs 1 credit per started 10 minutes, or 1 per started minute without captions; see [Credits & limits](https://cleanscript.ai/docs/credits.md). | Field | Type | Description | | --- | --- | --- | | `url` | string, required | A YouTube, TikTok or Instagram post link in any common form, or a bare YouTube video ID. | | `language` | string | The language to read in, such as `en` or `pt-BR`. Default: the language spoken in the post. Without captions in it, the transcript is translated and labelled. | | `include` | array of strings | Extra detail: `lines` (timed caption lines), `on_screen_text` (text shown in the video). Repeat the parameter in a query: `include=lines&include=on_screen_text`. | ```bash curl -G https://api.cleanscript.ai/v1/transcript \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "url=https://youtu.be/jwnez8HdN7E" ``` ```python result = client.transcript("https://youtu.be/jwnez8HdN7E") ``` ```typescript const result = await client.transcript("https://youtu.be/jwnez8HdN7E") ``` ```json Response {"post":{"platform":"youtube","type":"video","id":"jwnez8HdN7E","url":"https://www.youtube.com/watch?v=jwnez8HdN7E","title":"Microsoft’s new chip looks like science fiction…","description":"Sign up for CodeRabbit using FIRESHIP code, and get free CodeRabbit for 1-month https://bit.ly/41rLUxm …","author":{"name":"Fireship","handle":"@Fireship","url":"https://www.youtube.com/@Fireship"},"published_at":"2025-02-21","duration":259,"thumbnail_url":"https://i.ytimg.com/vi_webp/jwnez8HdN7E/maxresdefault.webp","sponsored":null,"sound":null,"stats":{"views":2331993,"likes":71026,"comments":3300,"shares":null,"saves":null}},"language":"en","source":"auto","text":"Out of nowhere, Microsoft just announced an impossible new quantum computing chip named Majorana 1, but it's not your average quantum chip. They claim to have created an entirely new state of matter. …\n\n…","sections":[{"title":"Microsoft's quantum breakthrough","start":2.19,"end":58.11,"url":"https://youtu.be/jwnez8HdN7E?t=2","paragraphs":[{"start":2.19,"end":58.11,"text":"Out of nowhere, Microsoft just announced an impossible new quantum computing chip named Majorana 1, but it's not your average quantum chip. They claim to have created an entirely new state of matter. …"}]}],"lines":null,"on_screen_text":null,"warnings":[]} ``` | Field | Type | Description | | --- | --- | --- | | `post` | object | The post. See [the post object](#the-post-object). | | `language` | string or null | The language of `text`; `null` when there is no speech. | | `source` | string | `creator`, `auto`, `transcribed`, `translated` or `none`; see [Transcripts](https://cleanscript.ai/docs/transcripts.md#where-the-words-come-from). | | `text` | string | The whole transcript, paragraphs separated by blank lines. | | `sections` | array | `{title, start, end, url, paragraphs}`; each paragraph is `{start, end, text}`. `title` is `null` for a short post. | | `lines` | array or null | `{start, end, text}` caption lines; only with `include: ["lines"]`. | | `on_screen_text` | array or null | `{start, end, slide, text}`. Filled when there is no speech, otherwise with `include: ["on_screen_text"]`. | | `warnings` | array | `{code, message}`; `code` is `translated`, `no_speech`, `uncorrected` (some lines couldn't be corrected) or `incomplete` (part of the post couldn't be read). Retry for all of it. | ## POST /v1/extract The fields you ask for, each with the quotes that support it. Costs 3 credits per started 10 minutes, the transcript included; see [Extract fields](https://cleanscript.ai/docs/extract.md). Send `prompt`, `preset` or `schema`; a `prompt` can go with either of the other two. | Field | Type | Description | | --- | --- | --- | | `url` | string, required | A YouTube, TikTok or Instagram post link, or a bare YouTube video ID. | | `prompt` | string | A question or instruction in plain words, up to 4,000 characters. Alone, it returns `{answer, points}`; with a schema or preset it adds guidance. | | `preset` | string | A published schema: `summary`, `mentions`, `ad-breakdown`, `claims`. | | `schema` | object | Your own JSON Schema, object at the root, up to 32 KB. | | `language` | string | The language to write values in. Default: the transcript's. Quotes stay word for word. | | `include` | array of strings | `transcript`: add the transcript the values come from. | ```bash curl https://api.cleanscript.ai/v1/extract \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ -H "Content-Type: application/json" \ -d '{"url": "https://youtu.be/jwnez8HdN7E", "preset": "mentions"}' ``` ```python result = client.extract("https://youtu.be/jwnez8HdN7E", preset="mentions") ``` ```typescript const result = await client.extract("https://youtu.be/jwnez8HdN7E", { preset: "mentions" }) ``` ```json Response {"post":{"platform":"youtube","type":"video","id":"jwnez8HdN7E","url":"https://www.youtube.com/watch?v=jwnez8HdN7E","title":"Microsoft’s new chip looks like science fiction…","description":"Sign up for CodeRabbit using FIRESHIP code, and get free CodeRabbit for 1-month https://bit.ly/41rLUxm …","author":{"name":"Fireship","handle":"@Fireship","url":"https://www.youtube.com/@Fireship"},"published_at":"2025-02-21","duration":259,"thumbnail_url":"https://i.ytimg.com/vi_webp/jwnez8HdN7E/maxresdefault.webp","sponsored":null,"sound":null,"stats":{"views":2331993,"likes":71026,"comments":3300,"shares":null,"saves":null}},"data":{"mentions":[{"name":"Microsoft","sentiment":"mixed","type":"company"}],"recommendations":[],"sponsors":[{"brand":"CodeRabbit","code":"FIRESHIP","link":"https://bit.ly/41rLUxm","offer":"One month free for teams","product":"AI code reviews"}]},"citations":{"mentions[0]":[{"quote":"If it turns out not to be your typical Microsoft BS—and that's a big if—it could be a breakthrough on par with the transistor.","source":"speech","start":20.5,"end":27.9,"url":"https://youtu.be/jwnez8HdN7E?t=20"}],"sponsors[0]":[{"quote":"It's 100% free for open source projects, but you can get one month free for your team using the code FIRESHIP with the link below.","source":"speech","start":249.7,"end":258.6,"url":"https://youtu.be/jwnez8HdN7E?t=249"},{"quote":"Sign up for CodeRabbit using FIRESHIP code, and get free CodeRabbit for 1-month https://bit.ly/41rLUxm","source":"description","start":null,"end":null,"url":null}]},"removed":[],"warnings":[],"transcript":null} ``` | Field | Type | Description | | --- | --- | --- | | `post` | object | The post. | | `data` | object | Your fields, matching the schema. Any value may be `null` when the post doesn't support it. | | `citations` | object | Quotes for each value by path (`offer`, `sponsors[0]`): `{quote, source, start, end, url}`, `source` being `speech`, `title`, `description` or `on_screen`. | | `removed` | array of strings | Paths of values dropped because no quote in the post supports them. | | `warnings` | array | `{code, message}`: `no_speech`, `uncorrected` or `incomplete`. | | `transcript` | object or null | The transcript the values come from, as `/v1/transcript` returns it without `lines`; only with `include: ["transcript"]`. | ## GET /v1/posts A creator's recent posts, newest first. Costs 1 credit; see [Creator posts](https://cleanscript.ai/docs/posts.md). | Field | Type | Description | | --- | --- | --- | | `author` | string, required | A profile or channel link in any form, or a handle such as `@Fireship` with `platform`. | | `platform` | string | `youtube`, `tiktok` or `instagram`; needed only when `author` is a handle. | | `limit` | integer | How many posts, 12 by default. At most 30 on YouTube, 35 on TikTok, 12 on Instagram. | | `since` | string | Only posts published on or after this date, `YYYY-MM-DD`. | ```bash curl -G https://api.cleanscript.ai/v1/posts \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \ --data-urlencode "author=https://www.youtube.com/@Fireship" ``` ```python listing = client.posts("https://www.youtube.com/@Fireship") ``` ```typescript const listing = await client.posts("https://www.youtube.com/@Fireship") ``` ```json Response {"author":{"platform":"youtube","name":"Fireship","handle":"@Fireship","url":"https://www.youtube.com/@Fireship","followers":4290000},"posts":[{"platform":"youtube","type":"video","id":"_5p1_TNSWqQ","url":"https://www.youtube.com/watch?v=_5p1_TNSWqQ","title":"PewDiePie is setting AI free... and OpenAI is furious","description":null,"author":{"name":"Fireship","handle":"@Fireship","url":"https://www.youtube.com/@Fireship"},"published_at":"2026-10-05","duration":346,"thumbnail_url":"https://i.ytimg.com/vi/_5p1_TNSWqQ/hq720_custom_2.jpg?sqp=CMjykdYG-oaymwEcCNAFEJQDSFXyq4qpAw4IARUAAIhCGAFwAcABBg==&rs=AOn4CLBewv6zstbI10M8d0myK-Ia1GS9sQ","sponsored":null,"sound":null,"stats":{"views":489000,"likes":null,"comments":null,"shares":null,"saves":null}}]} ``` | Field | Type | Description | | --- | --- | --- | | `author` | object | `{platform, name, handle, url, followers}`; `followers` is `null` when the platform doesn't show it. | | `posts` | array | [Post objects](#the-post-object), newest first. | ## GET /v1/account Your balance and rate limit. Free. ```bash curl https://api.cleanscript.ai/v1/account \ -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" ``` ```python account = client.account() ``` ```typescript const account = await client.account() ``` ```json Response {"email":"you@example.com","credits_remaining":940,"rate_limit_per_minute":10} ``` | Field | Type | Description | | --- | --- | --- | | `email` | string or null | The account's email. | | `credits_remaining` | integer | Credits left. | | `rate_limit_per_minute` | integer | Requests a minute this account may send. | ## The post object The same object in every response and in every item of a post list. Every field is always present; unknown values are `null`, never `0`. | Field | Type | Description | | --- | --- | --- | | `platform` | string | `youtube`, `tiktok` or `instagram`. | | `type` | string | `video`, `photo` or `carousel`. | | `id` | string | The platform's ID for the post. | | `url` | string | Link to the post. | | `title` | string or null | The YouTube title; `null` on TikTok and Instagram. | | `description` | string or null | The post's own text: the YouTube description, or the TikTok or Instagram caption. | | `author` | object | `{name, handle, url}`. | | `published_at` | string or null | Publication date, `YYYY-MM-DD`. | | `duration` | number or null | Length in seconds; `null` for photos. | | `thumbnail_url` | string or null | The YouTube thumbnail; `null` on TikTok and Instagram, whose image links expire. | | `sponsored` | boolean or null | The platform's paid-partnership or ad flag. | | `sound` | object or null | TikTok and Instagram audio: `{title, author, original}`, where `original` is true for the creator's own audio. `null` on YouTube. | | `stats` | object | `{views, likes, comments, shares, saves}` as the platform shows them. On Instagram, views come only in `/v1/posts` listings, for Reels; a single post's are `null`. |