# Extract fields

Ask a question, pick a preset or send your own JSON Schema; get JSON back with a quote and a moment for every value.

## Three ways to ask

`POST /v1/extract` takes the post `url` and one of these:

- `prompt`: a question or instruction in plain words, such as "What does it say about pricing?".
- `preset`: one of four ready-made schemas, `summary`, `mentions`, `ad-breakdown`, `claims`.
- `schema`: your own JSON Schema.

A `prompt` can also go with a `preset` or `schema`, as extra guidance: "focus on the sponsor segment".

```bash
curl https://api.cleanscript.ai/v1/extract \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://youtu.be/jwnez8HdN7E", "prompt": "What does it say about how the chip works?"}'
```

```python
result = client.extract("https://youtu.be/jwnez8HdN7E", prompt="What does it say about how the chip works?")
print(result.data["answer"])
```

```typescript
const result = await client.extract("https://youtu.be/jwnez8HdN7E", { prompt: "What does it say about how the chip works?" })
console.log(result.data.answer)
```

A `prompt` on its own returns `data` with an `answer` (`null` when the post doesn't address it) and the `points` it rests on (shortened):

```json Response
{
  "data": {
    "answer": "The chip uses Majorana zero modes at the ends of atomically engineered nanowires, which Microsoft says are resistant to quantum decoherence. It computes by braiding and fusing these particles, then measuring whether the wire contains an even or odd number of electrons; the wires are made into a “topoconductor” and can be linked together. The chip must be kept near absolute zero, and its claimed ability to scale to millions of qubits has not yet been demonstrated.",
    "points": [
      "The approach is based on Majorana fermions, described as particles that are their own antiparticles and highly resistant to decoherence."
    ]
  },
  "citations": {
    "points[0]": [
      {
        "quote": "It's based on the Majorana fermion, which is a subatomic particle that's also its own antiparticle.",
        "source": "speech",
        "start": 96.44,
        "end": 102.06,
        "url": "https://youtu.be/jwnez8HdN7E?t=96"
      }
    ]
  },
  "removed": [],
  "warnings": []
}
```

## What comes back

- `data`: your fields, matching the schema. Any value can be `null` when the post doesn't support it, so a schema never fails the request.
- `citations`: the quotes behind each value, keyed by its path.
- `removed`: paths of values we dropped because no quote in the post supports them.
- `warnings`: only when the result differs from the request: `incomplete` when part of the post couldn't be read, or `uncorrected`.
- `post`: the post, as on every route. Add `include: ["transcript"]` to get the transcript the values come from.

## Quotes

Every value comes with up to 3 quotes. A quote has the words as they appear in the post (`quote`; a spoken quote is widened to the whole sentences it is part of), where they appear (`source`: `speech`, `title`, `description` or `on_screen`), and, for speech and on-screen text in a video, its `start`, `end` and a `url` to that moment.

- Keys are readable paths: `offer`, `claims[1]`, `host.name`. An object in a list is quoted once as a whole: `sponsors[0]`.
- Neighbouring caption lines are joined into one quote.
- `null`, `[]`, `false` and `0` need no quote. A `true` always has one; a yes/no the post doesn't support comes back `false`.
- Values are checked against the post. A value no quote supports is removed and listed in `removed`.
- Values are written in the `language` you ask for (by default, the transcript's); quotes stay word for word.

## Presets

Each preset is a JSON Schema published on the site. Sending `"preset": "summary"` is the same as sending that file as `schema`; its field descriptions carry all the guidance.

| Preset | For | Fields |
| --- | --- | --- |
| [`summary`](https://cleanscript.ai/schemas/summary.json) | What a post says, for notes, newsletters and research | `summary`, `key_points`, `quotes`, `people` |
| [`mentions`](https://cleanscript.ai/schemas/mentions.json) | Brands, products and sponsors in a post, for brand monitoring and influencer marketing | `sponsors`, `mentions`, `recommendations` |
| [`ad-breakdown`](https://cleanscript.ai/schemas/ad-breakdown.json) | A strategist's breakdown of an ad or a promotional post, from what is said and the post's text | `hook`, `hook_type`, `angle`, `audience`, `pain_points`, `promise`, `product`, `offer`, `claims`, `proof`, `objections_handled`, `cta`, `structure`, `tone` |
| [`claims`](https://cleanscript.ai/schemas/claims.json) | What a post asserts, for research, fact-checking and brand safety | `claims`, `figures`, `sources` |

Presets only gain fields; a change that would break one gets a new name.

## Your own schema

Send any JSON Schema (draft 2020-12) with an object at the root, up to 32 KB; local `$ref`s work; recursive schemas and `pattern` don't. Describe each field in its `description`: that is what guides the answer.

```json Request
{
  "url": "https://www.youtube.com/watch?v=x7X9w_GIm1s",
  "language": "fr",
  "schema": {
    "properties": {
      "creator": {
        "description": "Who created the language.",
        "type": "string"
      },
      "release_year": {
        "type": "integer"
      },
      "use_cases": {
        "items": {
          "type": "string"
        },
        "type": "array"
      }
    },
    "required": [
      "creator",
      "release_year",
      "use_cases"
    ],
    "type": "object"
  }
}
```

## Write fields that work

- **Describe each field in plain words**, and say what it is not when two fields could overlap: "the deal offered; not the product".
- **Ask for the thing, not its time.** Every value's quote already carries `start`, `end` and a link.
- **Keep numbers as stated.** "three to six months" doesn't fit an integer. Use a string for the figure as said and another for what it measures.
- **Use an `enum` for anything you will group by** (type, sentiment, stage), with an `"other"` option.
- **Need exact words? Use the quote.** A description saying "word for word" is a request; the quote is the guarantee.
- **Prefer specific, quotable items** over whole-video judgements, and on long videos ask for "the 10 most important" rather than "every".
- **Ask about what is said or written.** Fields are filled from the speech, the title and the description (and the on-screen text when a post has no speech), not from images, music or sounds.
- **Give context in the root `description`**, such as "we track mentions of our brand and its competitors".

## Cost

3 credits per started 10 minutes of video, the transcript included. Without captions: the transcription, 1 credit per started minute, plus 2 per started 10 minutes. See [Credits & limits](https://cleanscript.ai/docs/credits.md).
