Skip to content

TranscriptFor developers, automation, agents

TikTok transcript API: captions, on-screen text, sounds

Get a TikTok's transcript from its link in one request: corrected captions or transcribed audio, on-screen text, translations, sound and ad flag.

3 min readBy CleanScriptMarkdown

On this page
  1. Make the request
  2. What comes back
  3. Text on screen
  4. Read it in another language
  5. The sound and the ad flag
  6. List a creator's recent TikToks
  7. What it costs
  8. Go further

To get a TikTok's transcript, send its link to /v1/transcript. You get the words as clean text with their times, the text shown on screen when you ask for it, and the post's details: views, likes, shares, saves, the sound it uses and whether TikTok marks it as an ad.

TikTok is in beta. Some requests are slower than on YouTube, and when TikTok doesn't answer you get source_unavailable with a Retry-After. A failed request costs nothing.

Make the request #

Create a key on the API keys page. New accounts get 100 free credits, no card.

curl
curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://www.tiktok.com/@washingtonpost/video/7692148602235309342" \
  -d include=lines
Python
# pip install cleanscript-ai
from cleanscript_ai import CleanScript

client = CleanScript()
result = client.transcript("https://www.tiktok.com/@washingtonpost/video/7692148602235309342", include=["lines"])
print(result.text)
JavaScript
// npm install cleanscript
import { CleanScript } from "cleanscript"

const client = new CleanScript()
const result = await client.transcript("https://www.tiktok.com/@washingtonpost/video/7692148602235309342", { include: ["lines"] })
console.log(result.text)

Video links, photo links and share links all work, with or without https://. Your server never calls TikTok.

What comes back #

The Washington Post's 77-second TikTok about the world's oldest bartender, shortened:

JSON
{
  "post": {
    "platform": "tiktok",
    "type": "video",
    "id": "7692148602235309342",
    "url": "https://www.tiktok.com/@washingtonpost/video/7692148602235309342",
    "title": null,
    "description": "For more than six decades, 103-year-old Irvin Koch has been pouring drinks for customers out of the basement of his home in Maryland. …",
    "author": { "name": "The Washington Post", "handle": "@washingtonpost", "url": "https://www.tiktok.com/@washingtonpost" },
    "published_at": "2026-10-04",
    "duration": 77.0,
    "sponsored": false,
    "sound": { "title": "original sound - The Washington Post", "author": "The Washington Post", "original": true },
    "stats": { "views": 1300000, "likes": 258200, "comments": 1352, "shares": 58600, "saves": 19238 }
  },
  "language": "en",
  "source": "auto",
  "text": "Hi, Mr. Cameron. How are you doing? Haha. Meet Irv, the world's oldest bartender. Right now I'm 103 years old. There we go. Irv runs a legal bar out of the basement of his home in Maryland, the last of its kind in the state. …",
  "lines": [
    { "start": 0.22, "end": 3.1, "text": "Hi, Mr. Cameron. How are you doing? Haha." },
    { "start": 3.1, "end": 6.26, "text": "Meet Irv, the world's oldest bartender." },
    { "start": 6.26, "end": 11.0, "text": "Right now I'm 103 years old. There we go." }
  ],
  "warnings": []
}
  • source: "auto" means the words come from TikTok's automatic captions, corrected: punctuation, capitals, misheard words, and names spelled as the post's description spells them. When the creator wrote their own captions, they come back as written (creator).
  • No captions isn't an error. We transcribe the audio and set source to transcribed, for posts up to 15 minutes.
  • lines are caption-sized lines with times in seconds, here because of include=lines. TikTok links can't open at a moment, so a section's url is the post link and start says where to look.
  • post is the same object on every route and every platform. A count TikTok doesn't show is null, never 0.

Text on screen #

Most TikToks print words over the video. Add include=on_screen_text and you get them with their times, next to the speech. On the same post:

JSON
"on_screen_text": [
  { "start": 0.0, "end": 6.46, "slide": null, "text": "At 103, the world's oldest bartender is still happy to pour you a drink" },
  { "start": 6.46, "end": 12.93, "slide": null, "text": "WORLD'S OLDEST BARTENDER" },
  { "start": 12.93, "end": 19.39, "slide": null, "text": "CASH ONLY" },
  { "start": 19.39, "end": 25.86, "slide": null, "text": "He and his brother bought the house in 1963." }
]

That adds 1 credit. When a post has no speech, you get the on-screen text without asking, and a no_speech warning. TikTok photo posts come back by slide. Duolingo's photo post with 1.4 million views:

JSON
"on_screen_text": [
  { "start": null, "end": null, "slide": 1, "text": "“you stupid”" },
  { "start": null, "end": null, "slide": 2, "text": "“no i’m not”" },
  { "start": null, "end": null, "slide": 3, "text": "“what’s 9 + 10?”" },
  { "start": null, "end": null, "slide": 4, "text": "“ … ”" },
  { "start": null, "end": null, "slide": 5, "text": "21" }
]

Read it in another language #

Add language to read a TikTok in your language. TikTok translates some posts itself, and when it has a translation in that language, you get it. A HelloFresh ad in French:

curl
curl -G https://api.cleanscript.ai/v1/transcript \
  -H "Authorization: Bearer $CLEANSCRIPT_API_KEY" \
  --data-urlencode "url=https://www.tiktok.com/@hellofresh/video/7359322279727189281" \
  -d language=fr
JSON
{
  "language": "fr",
  "source": "translated",
  "text": "Que se passerait-il si tu appelais ta mère?",
  "warnings": [{ "code": "translated", "message": "This is the platform's machine translation from English." }]
}

The original is the creator's own caption, "what would happen if you just called your mom up?". When TikTok has none, you get our translation, line by line with the times kept, for 1 credit more per started 10 minutes. The warning says which one you got: on the Washington Post video in Spanish, it reads "No Spanish captions, so this is our translation from English."

The sound and the ad flag #

Every TikTok's post says which sound it uses and whether TikTok marks it as an ad. The same HelloFresh post:

JSON
"sponsored": true,
"sound": { "title": "sonido original", "author": "- g 🦙", "original": true }
  • sponsored is TikTok's own ad flag.
  • sound.author is whose sound it is. Here HelloFresh uses another account's sound, which is common in trend videos.
  • original: false means licensed music. Songs aren't returned as speech: a post set to a song gives empty text, a no_speech warning, and its on-screen text.

List a creator's recent TikToks #

/v1/posts takes a profile link or a handle and returns the creator's latest posts, newest first, up to 35 per call. Each url goes straight into /v1/transcript or /v1/extract.

Python
listing = client.posts("https://www.tiktok.com/@duolingo", limit=20)
for post in listing.posts:
    print(post.published_at, post.stats.views, post.sponsored, post.url)

From Duolingo's list, shortened to one post:

JSON
{
  "author": { "platform": "tiktok", "name": "Duolingo", "handle": "@duolingo", "url": "https://www.tiktok.com/@duolingo", "followers": 18200000 },
  "posts": [
    {
      "type": "video",
      "url": "https://www.tiktok.com/@duolingo/video/7689488735649402126",
      "description": "i’m feeling TREMENDOUS (spike is not) 🎤: @Ray William Johnson @Your Favorite Martian …",
      "published_at": "2026-09-25",
      "duration": 8.0,
      "sponsored": true,
      "sound": { "title": "original sound", "author": "elmer.2847", "original": true },
      "stats": { "views": 3700000, "likes": 478000, "comments": 9952, "shares": 35800, "saves": 46122 }
    }
  ]
}

Add since=2026-10-01 to get only newer posts, for a daily check. See Creator posts.

What it costs #

  • A transcript costs 1 credit per started 10 minutes of video, so a TikTok with captions is 1 credit unless it runs past 10 minutes.
  • A TikTok without captions is transcribed for 1 credit per started minute instead: 1 credit for a 45-second post.
  • On-screen text next to speech adds 1 credit. A list of posts costs 1 credit.
  • The same request again within 6 hours is free, and so is any request that fails.

See pricing for packs.

Go further #

Spot something wrong? Tell us.