Skip to content

Creative breakdownFor marketers, developers

Find what a competitor's most viewed videos have in common

Rank a brand's recent TikToks by views, pull the hook and premise from the top and bottom five, and compare them. One Python script, run on Duolingo.

3 min readBy CleanScriptMarkdown

On this page
  1. The four steps
  2. The script
  3. What it printed
  4. Read the result
  5. Make it yours
  6. What it costs
  7. Go further

When a competitor's videos get ten times the views of yours, the question is what their best ones share. You can answer it in four steps: list their recent posts, rank them by views, read the same fields from the top and the bottom, and compare.

This won't predict which video will take off next. It shows you what the posts that already did well have in common, so you know what to test.

The four steps #

  1. List the posts. /v1/posts returns a creator's latest posts with views, length, the sound used and the ad flag, for 1 credit. Up to 35 per call on TikTok.
  2. Rank them. Sort by stats.views and take the top five and the bottom five. The bottom group is what makes the comparison mean something.
  3. Read the same fields from each. /v1/extract with your own schema returns the hook, the premise and whatever else you want to compare, each value with the words it came from.
  4. Compare. Put both groups in one table and count what differs.

The script #

Python
# pip install cleanscript-ai
from statistics import median

from cleanscript_ai import CleanScript

CREATOR = "https://www.tiktok.com/@duolingo"

SCHEMA = {
    "type": "object",
    "description": "We compare a brand's most and least viewed short videos to see what the top ones share.",
    "properties": {
        "hook": {"type": ["string", "null"], "description": "The first words the viewer hears or reads, word for word."},
        "hook_type": {
            "type": ["string", "null"],
            "enum": ["question", "bold_claim", "problem", "story", "relatable", "curiosity", "news", "offer", "other", None],
            "description": "The hook's technique. relatable: a situation the viewer recognises from their own life.",
        },
        "premise": {"type": ["string", "null"], "description": "The joke or idea of the post, in one short sentence."},
        "kind": {
            "type": "string",
            "enum": ["language_joke", "lesson", "promotion", "announcement", "other"],
            "description": "language_joke: humour that plays on a word or phrase in another language. lesson: teaches something. promotion: promotes a product, partner or event.",
        },
        "languages": {"type": "array", "items": {"type": "string"}, "description": "Languages other than English that the post uses or teaches."},
        "partners": {"type": "array", "items": {"type": "string"}, "description": "Other brands, games or creators the post names or tags."},
    },
    "required": ["hook", "hook_type", "premise", "kind", "languages", "partners"],
}

client = CleanScript()
listing = client.posts(CREATOR, limit=20)
creator = listing.author.name

ranked = sorted((p for p in listing.posts if p.stats.views), key=lambda p: p.stats.views, reverse=True)
groups = {"top": ranked[:5], "bottom": ranked[-5:]}

rows = []
for group, posts in groups.items():
    for post in posts:
        data = client.extract(post.url, schema=SCHEMA, language="en").data
        borrowed = bool(post.sound and post.sound.author != creator)
        rows.append({"group": group, "post": post, "data": data, "borrowed": borrowed})

print(f"{'group':<7}{'views':>10}{'secs':>6}  {'sound':<9}{'kind':<14}{'hook_type':<11}hook")
for row in rows:
    post, data = row["post"], row["data"]
    secs = f"{post.duration:.0f}" if post.duration else "photo"
    sound = "borrowed" if row["borrowed"] else "own"
    hook = (data["hook"] or "")[:44]
    print(f"{row['group']:<7}{post.stats.views:>10,}{secs:>6}  {sound:<9}{data['kind'] or '-':<14}{data['hook_type'] or '-':<11}{hook}")

print()
for group in groups:
    chosen = [row for row in rows if row["group"] == group]
    lengths = [row["post"].duration for row in chosen if row["post"].duration]
    print(
        f"{group}: median {median(lengths):.0f}s, "
        f"{sum(row['borrowed'] for row in chosen)}/{len(chosen)} borrowed sounds, "
        f"{sum(row['data']['kind'] == 'promotion' for row in chosen)}/{len(chosen)} promotions, "
        f"{sum(bool(row['data']['languages']) for row in chosen)}/{len(chosen)} use another language, "
        f"{sum(bool(row['post'].sponsored) for row in chosen)}/{len(chosen)} marked as ads"
    )

Some of the comparison costs nothing beyond the list: length, views, the sound and the ad flag come with every post. The schema adds what only the words can tell you. language="en" writes the values in English even when a post speaks Spanish or Japanese; the hook stays word for word.

What it printed #

Run on Duolingo's 20 latest TikToks:

Text
group       views  secs  sound    kind          hook_type  hook
top     3,700,000     8  borrowed language_joke bold_claim I’m feeling TREMENDOUS
top     2,000,000    12  borrowed other         question   ¿No jugaste Roblox conmigo?
top     1,400,000 photo  borrowed other         other      you stupid
top     1,200,000    21  borrowed -             other      헬로우 에브리니언.
top     1,200,000     7  borrowed lesson        other      Hey, don't say that. Really freaking easy.
bottom    200,400    13  borrowed promotion     other      Pizza Nizar Saint-Gilles, mahbabi koum, c'es
bottom    187,200     9  own      promotion     -          It's a beautiful day.
bottom    104,700     6  borrowed announcement  other      Get ready for work without me.
bottom    100,800 photo  borrowed language_joke other      TACO BELL 101
bottom     86,600     7  own      language_joke question   Taco?

top: median 10s, 5/5 borrowed sounds, 0/5 promotions, 2/5 use another language, 3/5 marked as ads
bottom: median 8s, 3/5 borrowed sounds, 2/5 promotions, 2/5 use another language, 3/5 marked as ads

Read the result #

What the top five share. All five use a sound from another account, and none of them sells anything. They are jokes set to another account's sound: a Spanish line about Roblox, a Japanese exchange about speaking English, a "9 + 10" photo meme.

What sets the bottom apart. Both posts on Duolingo's own sound are here, and three of the five belong to its Taco Bell partnership ("new job who dis @tacobell", "TACO BELL 101", "Taco?"). Partner posts got a fraction of the views of the jokes.

What doesn't separate them. Length is close (a median of 10 seconds against 8), as is the share of posts in another language. TikTok marks three of each five as ads, so the ad flag doesn't explain the gap.

Check the quotes before you conclude. The bottom group's first "promotion" is a "translate this, wrong answers only" post. Its citation shows why it was filed that way:

JSON
"kind": [{ "quote": "La meilleure pizza, wallah, ingrédients frais. Dégustez le bol.", "source": "speech", "start": 6.16, "end": 13.0, "url": "https://www.tiktok.com/@duolingo/video/7691397360794029343" }]

The speech is a pizza shop's ad, used as the borrowed sound. Duolingo isn't selling pizza, so count it as a joke. And when a value has no quote to back it, it is removed: the fourth top video shows - because kind came back null and is listed in removed.

This is one brand and 20 posts over four weeks. The finding is a lead: "jokes on borrowed sounds beat partner posts". Test it on your own account before you plan around it.

Make it yours #

  • Other questions: change the schema. Add a cta field to see which posts ask for anything, or an audience field. Fields are filled from what is said and written in the caption (and the text on screen when a post has no speech), not from the images, so ask about words. See how to write fields that work.
  • Paid ads: swap schema=SCHEMA for preset="ad-breakdown". It returns the hook, offer, claims, call to action and the ad's stages in order. See the ad breakdown guide.
  • Instagram: use a profile link such as https://www.instagram.com/natgeo/. Photos and carousels have no views, so rank by stats.likes. A call returns up to 12 posts.
  • YouTube: use a channel link. The list holds videos, not Shorts, with view counts rounded.
  • A bigger sample: TikTok returns up to 35 posts per call. Run the script every few weeks with since and keep the rows, and the comparison gets stronger with each run.

What it costs #

The list costs 1 credit. Each short video costs 3 credits, so this run was ten times that plus the list. A post you already transcribed costs only the fields, and running the same script again within 6 hours is free. See pricing.

Go further #

Spot something wrong? Tell us.