Blog / Endpoint guides

Scrape Instagram Comments With Cursor Pagination

By the InScrape API team · Published 2026-09-02 · 7 min read

Speech bubbles arranged along a stepped pagination path

Comments are the only part of an Instagram post that is written by the audience rather than the brand. Likes tell you how many people reacted; comments tell you what they said, what they wanted, and whether they were a person at all.

They are also the endpoint most people get wrong, because the first page looks like success. Instagram hands a logged-out visitor one page of ranked comments, and a naive script stops there and reports it as "the comments". On a post with four thousand comments that page is a rounding error, and a ranked rounding error at that — the slice most biased towards what the platform already chose to promote.

This is how to take the whole thread.

One request, by URL or shortcode

The url parameter accepts either a full post URL or the bare shortcode. Both resolve to the same post:

curl "https://api.socialscrape.dev/v1/instagram/comments?url=https://www.instagram.com/p/C8QltIdyBGH/" \
  -H "x-api-key: $INSCRAPE_KEY"

curl "https://api.socialscrape.dev/v1/instagram/comments?url=C8QltIdyBGH" \
  -H "x-api-key: $INSCRAPE_KEY"

Reel URLs work too: /reel/{shortcode} is the same object with a different path segment, so nothing needs rewriting before you send it.

One thing worth knowing before you build a pipeline on this. Pick one post identifier format at the point where the URL enters your system and keep it. Mixing full URLs and bare shortcodes makes your own logs and dedupe rules harder to reason about.

The response is the standard envelope:

{
  "success": true,
  "credits_charged": 1,
  "credits_remaining": 14872,
  "processing_time_ms": 1842,
  "requested_at": "2026-09-13T14:32:18Z",
  "query": { "url": "C8QltIdyBGH" },
  "data": {
    "comments": [
      {
        "id": "17984412330912345",
        "text": "how much is shipping to the UK? 😩",
        "created_at": "2026-08-30T14:02:11Z",
        "like_count": 12,
        "reply_count": 1,
        "author": {
          "handle": "mara.builds",
          "full_name": "Mara",
          "profile_pic_url": "https://..."
        }
      }
    ],
    "next_cursor": "QVFEZ..."
  }
}

One credit per page. next_cursor is null on the last page, and that is the only reliable stop condition — comment pages are not a fixed size, and an empty comments array can appear mid-thread.

Draining the cursor

The loop is four lines and one guard. The guard is the part that matters: a thread has no depth you know in advance, so a loop without a cap keeps spending credits for as long as the post keeps handing back cursors.

import os
import requests

BASE = "https://api.socialscrape.dev/v1"
HEADERS = {"x-api-key": os.environ["INSCRAPE_KEY"]}


def comments(url: str, max_pages: int = 20, **extra):
    cursor = None
    for page in range(max_pages):
        params = {"url": url, **extra}
        if cursor:
            params["cursor"] = cursor

        res = requests.get(f"{BASE}/instagram/comments", params=params,
                           headers=HEADERS, timeout=30)
        body = res.json()

        if not res.ok:
            code = body.get("error", {}).get("code")
            # Missing resources and invalid parameters are not transient failures.
            if code in {"ENDPOINT_NOT_FOUND", "INVALID_PARAMS"}:
                raise LookupError(code)
            raise RuntimeError(code)

        if "profile" in body["data"] and body["data"]["profile"] is None:
            return
        yield from body["data"]["comments"]

        cursor = body["data"].get("next_cursor")
        if not cursor:
            return

    raise RuntimeError(f"Hit the {max_pages}-page cap with a cursor still open")
thread = list(comments("https://www.instagram.com/p/C8QltIdyBGH/", max_pages=40))
print(len(thread), "comments,", len({c["author"]["handle"] for c in thread}), "distinct authors")

Raising at the cap rather than returning quietly is deliberate. A silent truncation looks identical to a short thread in your data, and you will not notice until someone asks why a viral post has exactly 800 comments.

The Node version, as an async generator so the consumer decides when to stop:

const BASE = "https://api.socialscrape.dev/v1";

async function* comments(url, { maxPages = 20, ...extra } = {}) {
  let cursor = null;

  for (let page = 0; page < maxPages; page++) {
    const params = new URLSearchParams({ url, ...extra });
    if (cursor) params.set("cursor", cursor);

    const res = await fetch(`${BASE}/instagram/comments?${params}`, {
      headers: { "x-api-key": process.env.INSCRAPE_KEY },
    });
    const body = await res.json();

    if (!res.ok) throw new Error(body.error?.code ?? res.status);

    yield* body.data.comments;

    cursor = body.data.next_cursor;
    if (!cursor) return;
  }

  throw new Error(`Hit the ${maxPages}-page cap with a cursor still open`);
}

const thread = [];
for await (const c of comments("C8QltIdyBGH", { maxPages: 40 })) {
  thread.push(c);
  if (thread.length >= 2000) break; // consumer-side budget
}

Breaking out of the for await stops the generator, so no further pages are fetched and no further credits are spent. That is the useful property of writing it this way rather than collecting everything into an array first.

Pick the refresh schedule by post age

On an eighteen-month-old post, the thread is barely moving. On a post published forty minutes ago, comments can change minute by minute. Use the post age to decide how often your job should pull the thread.

For a launch post, a five-minute or thirty-minute schedule can make sense. For older posts, hourly or daily is usually enough:

curl "https://api.socialscrape.dev/v1/instagram/comments?url=C8QltIdyBGH" \
  -H "x-api-key: $INSCRAPE_KEY"
Post age Sensible refresh Why
Under 6 hours Every 5 minutes The thread is still forming
6-48 hours Every 30 minutes Slowing down, but a support complaint still needs picking up the same day
Over 48 hours Hourly or daily Old threads rarely need high-frequency refreshes

One warning worth stating plainly: a fresh pull is a fresh charge. Polling a live post every five minutes for a day is 288 credits on that post alone. Poll the posts that matter, not the feed.

What the comments are actually for

Collecting them is the easy half. Four things people build on the output, and what each one honestly requires:

Sentiment. We return text; we do not return a sentiment score, and nobody should sell you one bundled into a scraping API. You run the classifier — a model call, a lexicon, whatever fits your budget. The one thing worth knowing before you start is that Instagram comment text is emoji-heavy and multilingual, and a lexicon tuned on product reviews will misread both. Strip nothing before classification: the emoji frequently carries the polarity that the words do not.

Demand signals. The highest-value comments on a commercial post are questions, and questions have shape: a question mark, a size or colour noun, a country name, "restock", "link". Bucket them and count. A post with forty "ship to Canada?" comments is a market research result that cost you three credits to obtain.

Spam and bot detection. The response gives you enough to do this without a second endpoint. Duplicate text across distinct authors, comments that are nothing but emoji plus a mention, and clusters of created_at values seconds apart are all visible in one thread. What is not visible is the commenter's own follower count or post history — that needs a profile call per handle, at one credit each, so filter on the free signals first and enrich only the survivors.

Buying intent. Same mechanics as demand signals, different threshold. Comments naming a price, asking about availability, or tagging a second person are the ones worth routing to a human. Note that you are reading intent from a public comment, not from a purchase — the correlation is yours to validate against your own conversion data, and it varies enormously by category.

from collections import Counter

BUYING = ("price", "how much", "where can i buy", "link", "restock", "ship to")

intent = [c for c in thread if any(k in c["text"].lower() for k in BUYING)]
dupes = Counter(c["text"].strip() for c in thread)
bots = [c for c in thread if dupes[c["text"].strip()] > 3]

That keyword list is a starting point and an English-only one. Treat it as something you extend against your own threads, not a finished classifier.

What is not in the response

Being precise here saves refunds, so:

  • No email addresses or phone numbers. Not on the comment, not on the author, and not through any enrichment step. If a provider offers you commenter emails, ask them where those came from.
  • No private replies. A "private reply" on Instagram is a DM sent in response to a comment. DMs are not public data and we cannot read them, in any direction.
  • No hidden or filtered comments. Instagram's offensive-comment filter, an account's custom keyword filter and manually hidden comments all remove a comment from what a logged-out visitor renders. We read what that visitor sees, so those comments are not in the response and never will be. If you are auditing your own account's hidden queue, that is the Graph API on an account you own, not this.
  • No writes. You cannot post, reply to, like, hide or delete a comment through this API. It is read-only by design, which is also why it needs no login and puts no account at risk.
  • Ranked order, not chronological. The ordering matches what Instagram serves publicly. Sort by created_at yourself if you need a timeline.
  • Confirmed private accounts return 200 with available profile details, billed at the endpoint rate. A post on a private account has no public comment thread to read.
  • Deleted comments vanish. Each pull is a snapshot. If you need a durable record of what was said, store your own copy — re-fetching a week later will not bring back what the author removed.

Where to go next

Free tier is 100 credits, no card required. Billing depends on whether scraping actually used an Instagram or H API resource, not on the HTTP status.

Try it with 100 free credits.

No credit card, credits never expire, and failed requests are not charged.