Skip to content

Your fieldsFor developers, automation, newsletters & research

Extract structured data from a video with a JSON Schema

Describe the fields you want as a JSON Schema, send a YouTube, TikTok or Instagram link, and get JSON where every value has its quote and timestamp.

3 min readBy CleanScriptMarkdown

On this page
  1. A first request
  2. What comes back
  3. Write fields that work
  4. Presets
  5. Plain questions
  6. Quote a video accurately
  7. Cost
  8. Next steps

To pull structured data out of a video, describe the fields you want as a JSON Schema and send it with the link to POST /v1/extract. You get an object that matches your schema and, for every value, the words in the post that back it and a link to that moment.

A first request #

The example is The Economist's Instagram reel on Olympic medals. YouTube and TikTok links work the same way.

Python
# pip install cleanscript-ai
from cleanscript_ai import CleanScript

client = CleanScript()
result = client.extract(
    "https://www.instagram.com/reel/C-TCp_KhMr4/",
    schema={
        "type": "object",
        "properties": {
            "top_country": {"type": "string", "description": "The country with the most Olympic medals, as the post states it."},
            "medal_forecast": {"type": "integer", "description": "The number of medals the post says it is predicted to win in Paris."},
            "findings": {"type": "array", "items": {"type": "string"}, "description": "Each statistic the post gives, one per item, with its number."},
        },
        "required": ["top_country", "medal_forecast", "findings"],
    },
)
print(result.data["medal_forecast"])
print(result.citations["medal_forecast"][0].quote)

Over HTTP, send the same url and schema as a JSON body:

curl
curl https://api.cleanscript.ai/v1/extract \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.instagram.com/reel/C-TCp_KhMr4/", "schema": {…}}'

What comes back #

JSON
{
  "data": {
    "top_country": "America",
    "medal_forecast": 115,
    "findings": [
      "Since 2000, America won 11% of all Summer Olympic medals.",
      "Since 2000, China won 8% of all Summer Olympic medals.",
      "Adjusted for GDP, Jamaica ranks first and America ranks 50th.",
      "For every $100 billion of GDP, Jamaica wins 69 medals."
    ]
  },
  "citations": {
    "medal_forecast": [{
      "quote": "It's predicted that the US will take home 115 of about 1,000 medals.",
      "source": "speech",
      "start": 14.85,
      "end": 22.63,
      "url": "https://www.instagram.com/reel/C-TCp_KhMr4/"
    }]
  },
  "removed": [],
  "warnings": []
}

Shortened: the real findings has eight items, each with its own citation.

  • data matches your schema. Any value can be null when the post doesn't say, so a schema never fails the request.
  • citations holds the quotes for each value, keyed by its path: medal_forecast, findings[2]. source says where the quote is: speech, title, description or on_screen. On YouTube, url opens the video at that moment; on TikTok and Instagram it is the post link, and start says where to look.
  • removed lists values we dropped because no quote in the post supports them.

null, [], false and 0 need no quote. A true always has one.

Write fields that work #

  • Describe each field in plain words. The description is the instruction.
  • Ask for what the post states. "The price the speaker quotes" works; "is this a good deal?" doesn't.
  • Don't add placeholders like "N/A". A missing value is null.
  • Prefer lists of short items to one long string, and an enum when the answers are a fixed set.
  • Don't ask for timestamps. Every quote already carries start, end and a link.

A prompt next to the schema adds guidance, such as "focus on the sponsor segment". Schemas can be up to 32 KB.

Presets #

For common jobs, send a preset instead of a schema:

Preset Fields
summary summary, key_points, quotes, people
mentions sponsors (brand, offer, code, link), mentions (name, type, sentiment), recommendations
ad-breakdown hook, angle, audience, offer, claims, cta and more (see the guide)
claims claims (with who makes them), figures, sources

Each preset is a JSON Schema published on the site, such as /schemas/mentions.json. Copy one as a starting point for your own.

Run on Duolingo's TikTok The truth., mentions returns:

JSON
"mentions": [
  { "name": "Duolingo", "type": "app", "sentiment": "positive" },
  { "name": "TikTok", "type": "app", "sentiment": "neutral" },
  { "name": "Instagram", "type": "app", "sentiment": "neutral" },
  { "name": "LinkedIn", "type": "app", "sentiment": "negative" }
]

Each entry has its quote: mentions[1], TikTok, cites "I guess you could have your TikTok back for now." at 114 seconds. sponsors is empty because the post has none.

Plain questions #

For one question, send a prompt alone. You get data.answer and data.points, with citations. The same reel with "prompt": "Which country jumps to the top once medals are adjusted for GDP, and why?":

JSON
{
  "data": {
    "answer": "Jamaica jumps to the top because it earns 69 medals for every $100 billion of GDP when medal totals are adjusted for economic size.",
    "points": ["After adjusting for GDP, Jamaica ranks first and earns 69 medals per $100 billion of GDP."]
  },
  "citations": {
    "answer": [{ "quote": "For every 100 billion dollars of GDP, Jamaica nets 69 medals.", "source": "speech", "start": 47.31, "end": 60.0, "url": "https://www.instagram.com/reel/C-TCp_KhMr4/" }]
  }
}

If the post doesn't address the question, answer is null.

Quote a video accurately #

A citation has what you need to quote. From Steve Jobs' Stanford address, the summary preset's quote for the closing line:

JSON
{
  "quote": "Stay Hungry. Stay Foolish.",
  "source": "speech",
  "start": 868.68,
  "end": 871.61,
  "url": "https://youtu.be/UF8uR6Z6KLc?t=868"
}

Use the words as they are and link the url. APA and MLA both want a time stamp for a quote from a video: 868 seconds is 14:28. The quote keeps the captions' punctuation, so if a style needs more, play the link and check.

Cost #

An extraction costs 3 credits per started 10 minutes of video, the transcript included. See pricing.

Next steps #

Sources

  1. Direct quotation of material without page numbers - APA Style
  2. Time stamps for videos - MLA Style Center

Spot something wrong? Tell us.