# Instagram Reels transcript: the words and on-screen text

> Get an Instagram Reel's transcript from its link, with the text shown on screen. Photo and carousel posts give their text slide by slide.

Updated 2026-10-06. HTML: /blog/instagram-reels-transcript

To get the transcript of an Instagram Reel, paste its link into the [playground](/playground) or send it to `/v1/transcript`. You get what is said as clean text with times. A Reel told in captions on screen, or a carousel of text slides, gives you that text instead.

Instagram is in beta. Some requests are slower than on YouTube, and when Instagram doesn't answer you get `source_unavailable` with a `Retry-After`. A failed request costs nothing.

## Make the request

```bash
curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://www.instagram.com/reel/DeCPVhERA8e/" \
  -d include=lines
```

```python
# pip install cleanscript-ai
from cleanscript_ai import CleanScript

client = CleanScript()
result = client.transcript("https://www.instagram.com/reel/DeCPVhERA8e/", include=["lines"])
print(result.text)
```

`/reel/`, `/p/` and `/tv/` links all work, with or without `https://`. Only public posts can be read.

## A Reel with speech

Astronaut Jessica Meir's [Reel from the space station](https://www.instagram.com/reel/DeCPVhERA8e/), which also appears on NASA's profile, shortened:

```json
{
  "post": {
    "platform": "instagram",
    "type": "video",
    "id": "DeCPVhERA8e",
    "description": "Room with a view. Check out my new digs as I camp out in Dragon until our departure. …",
    "author": { "name": "Jessica Meir", "handle": "@astro_jessica", "url": "https://www.instagram.com/astro_jessica/" },
    "published_at": "2026-10-03",
    "duration": 27.56,
    "sponsored": false,
    "sound": { "title": "Original audio", "author": "astro_jessica", "original": true },
    "stats": { "views": null, "likes": 173287, "comments": 1816, "shares": null, "saves": null }
  },
  "language": "en",
  "source": "transcribed",
  "text": "Good morning from my Dragon lair. Now that Crew-13 has arrived, I've given up my crew quarters and I'm camping out in my Dragon vehicle. Here's my view right next to my sleeping bag outside my window here: the Earth and the Crew-13 Dragon. A gorgeous view. Reminds me of my best piece of advice to them: Remember to look out the window and realize where you are. …",
  "lines": [
    { "start": 0.16, "end": 2.2, "text": "Good morning from my Dragon lair." },
    { "start": 2.6, "end": 4.54, "text": "Now that Crew-13 has arrived," },
    { "start": 4.64, "end": 7.88, "text": "I've given up my crew quarters and I'm camping out in my Dragon vehicle." }
  ],
  "warnings": []
}
```

- **`source: "transcribed"`**: the Reel had no captions to use, so we transcribed its audio and corrected it, with names spelled as the post's description spells them. Reels up to 15 minutes are transcribed.
- **`post.author`** is the account that posted the Reel, here Jessica Meir, even when you found it on NASA's profile.
- **`stats`** has what Instagram shows: likes and comments. Shares and saves are `null`, and so is `views` on a single post.
- Instagram links can't open at a moment, so each section's `url` is the post link and `start` says where to look.

## A Reel without speech

Many Reels tell the story in text on screen over music or natural sound. When there is no speech, you get that text automatically, with its times. National Geographic's [black rain frog Reel](https://www.instagram.com/reel/DeCbRyHgmlu/):

```json
{
  "source": "none",
  "text": "",
  "on_screen_text": [
    { "start": 0.0, "end": 6.09, "slide": null, "text": "A rarely seen black rain frog emerges after six months underground." },
    { "start": 6.09, "end": 12.18, "slide": null, "text": "The frog surfaces in South Africa for one important reason: breeding season." },
    { "start": 12.18, "end": 15.22, "slide": null, "text": "Built for burrowing, he can’t hop or swim." },
    { "start": 15.22, "end": 21.31, "slide": null, "text": "Instead, the male uses a high-pitched call to attract a mate through the forest." },
    { "start": 21.31, "end": 24.36, "slide": null, "text": "And it works." },
    { "start": 24.36, "end": 30.45, "slide": null, "text": "But when she gets too close, he sounds an alarm call." }
  ],
  "warnings": [{ "code": "no_speech", "message": "This post has no speech." }]
}
```

Shortened: the full list also has the title cards, "NATIONAL GEOGRAPHIC", "EARTH'S WILD HOME", "NOW STREAMING". Text shown again later is listed again.

For a Reel with speech, add `include=on_screen_text` to get both. That adds 1 credit.

## Photo and carousel posts

A photo or carousel post has no speech, so the answer is the text on each image, by `slide`. Duolingo's [carousel](https://www.instagram.com/p/Ddrs0etFMAa/) captioned "nein out of ten, would learn again", with `post` shortened:

```json
{
  "post": { "type": "carousel", "id": "Ddrs0etFMAa", "duration": null, "sound": null },
  "source": "none",
  "on_screen_text": [
    { "start": null, "end": null, "slide": 1, "text": "germans when they say:" },
    { "start": null, "end": null, "slide": 1, "text": "“thank you”" },
    { "start": null, "end": null, "slide": 2, "text": "“3”" },
    { "start": null, "end": null, "slide": 3, "text": "germans when they say: “pants”" },
    { "start": null, "end": null, "slide": 4, "text": "germans when they say: “10”" },
    { "start": null, "end": null, "slide": 5, "text": "“bright”" },
    { "start": null, "end": null, "slide": 6, "text": "“11”" }
  ],
  "warnings": [{ "code": "no_speech", "message": "This is a carousel post, so it has no speech; on_screen_text holds any text shown in it." }]
}
```

Slides are numbered from 1. A slide with no text has no entry, so a carousel of photos returns little: National Geographic's 2,000-rhino carousel gives only "AFRICAN PARKS" and "RHINO REWILD", from its first slide. The story there is in `post.description`.

## List a profile's recent posts

`/v1/posts` takes a profile link or a handle and returns the latest posts, newest first, up to 12 per call: Reels, photos and carousels.

```python
listing = client.posts("https://www.instagram.com/natgeo/", limit=12)
for post in listing.posts:
    print(post.published_at, post.type, post.stats.likes, post.sponsored, post.url)
```

Two things to know when you compare posts:

- **Views:** in a list, Reels usually carry their `views`; photos and carousels have none. Compare Instagram posts by `likes` when you mix types.
- **`sponsored`** is Instagram's paid-partnership flag. In National Geographic's list it is `true` on three posts, while others that open with "Presented by @Rolex." have `false`. Read `description` too.

## What it costs

- A Reel with captions costs 1 credit: transcripts are priced per started 10 minutes of video.
- A Reel without captions is transcribed for 1 credit per started minute: 1 credit for the 28-second Reel above, 2 credits for a 90-second one.
- A Reel without speech, a photo or a carousel costs the transcript price for its on-screen text. With speech, on-screen text adds 1 credit.
- A list of posts costs 1 credit, and the same list again within 10 minutes is free. Asking again for the same transcript within 6 hours is free, and so is any request that fails.

## Go further

- Ask for fields instead of the whole text: [extract data with a JSON Schema](/blog/json-schema-fields-from-video).
- Find what a brand's best posts share: [compare a creator's top videos](/blog/competitor-top-videos-analysis).
- The same for TikTok, with its translations and sounds: [TikTok transcript API](/blog/tiktok-transcript-api).
