Blog / Developer tutorials

Build a Production Instagram Data Pipeline in Python

By the InScrape API team · Published 2026-08-12 · 4 min read

Code braces surrounding a glass data cube connected to a modular pipeline

Most Python Instagram tutorials show a single requests.get and stop. That works until you paginate, hit a private account, or run it against a thousand handles and discover you have no idea what it cost. If you only need copy-paste request examples, use the focused Instagram API Python reference; this guide is about the surrounding production pipeline.

This is the version that survives contact with production.

Setup

pip install requests
export INSCRAPE_KEY=sk_your_key

No SDK. One header, one base URL, and the standard library is enough — which is the point of a REST API with a consistent contract.

A client worth reusing

import os
import time
import requests

BASE = "https://api.socialscrape.dev/v1"


class InScrape:
    def __init__(self, key: str | None = None):
        self.session = requests.Session()
        self.session.headers["x-api-key"] = key or os.environ["INSCRAPE_KEY"]

    def get(self, path: str, **params):
        for attempt in range(3):
            res = self.session.get(f"{BASE}{path}", params=params, timeout=30)
            body = res.json()

            if res.ok:
                return body

            code = body.get("error", {}).get("code")

            # These are facts about the target, not transient failures.
            if code in {"ENDPOINT_NOT_FOUND", "INVALID_PARAMS"}:
                raise LookupError(code)

            # Retrying will not conjure credits.
            if code == "INSUFFICIENT_CREDITS":
                raise RuntimeError("Out of credits")

            # Everything else is worth one more try.
            time.sleep(2 ** attempt)

        raise RuntimeError(f"Failed after 3 attempts: {path}")

The retry policy is the important part. Validation and authentication failures need a corrected request. Confirmed private or missing Profile accounts return 200 with available profile details or data.profile=null. These completed lookups are billed, so record the outcome instead of retrying it.

Profiles

ig = InScrape()

body = ig.get("/instagram/profile", handle="natgeo")
data = body["data"]
if body.get("data", {}).get("profile") is None:
    print(body["data"]["profile"], "charged:", body["credits_charged"])
else:
    profile = data["profile"]
    print(profile["follower_count"], profile["category"])
print("requested:", body["requested_at"], "credits left:", body["credits_remaining"])

Every response carries the cost on it. Log credits_charged alongside your own request logs and you will never be surprised by an invoice.

Pagination, done once

Every list endpoint uses the same cursor. Write the loop once and reuse it everywhere:

from typing import Iterator


def paginate(ig: InScrape, path: str, key: str, max_pages: int = 10, **params) -> Iterator[dict]:
    cursor = None
    for _ in range(max_pages):
        body = ig.get(path, **params, **({"cursor": cursor} if cursor else {}))
        data = body["data"]

        yield from data.get(key, [])

        cursor = data.get("next_cursor")
        if not cursor:
            return
posts = list(paginate(ig, "/instagram/posts", "posts", max_pages=3, handle="natgeo"))
comments = list(paginate(ig, "/instagram/comments", "comments",
                         url="https://www.instagram.com/p/C8QltIdyBGH/"))

Always pass max_pages. A comment thread on a viral post can go deeper than your budget, and one credit per page adds up quickly. An unbounded pagination loop is the most expensive bug in this category.

Engagement rate, computed properly

Two endpoints, one number:

def engagement_rate(ig: InScrape, handle: str, sample: int = 12) -> float:
    profile = ig.get("/instagram/profile", handle=handle)["data"]
    posts = list(paginate(ig, "/instagram/posts", "posts",
                          max_pages=1, handle=handle, limit=sample))

    if not posts or not profile["follower_count"]:
        return 0.0

    interactions = sum(p["like_count"] + p["comment_count"] for p in posts)
    return interactions / len(posts) / profile["follower_count"] * 100

Two credits per creator. Note that this is your formula — Instagram does not publish an engagement rate, so document your denominator or your customers will compare it against someone else's and think one of you is broken.

Concurrency without a queue

There is no published rate limit, so throughput is bounded by your own connection pool. A thread pool covers most batch jobs:

from concurrent.futures import ThreadPoolExecutor


def fetch_many(handles: list[str], workers: int = 16) -> dict[str, dict | None]:
    ig = InScrape()

    def one(handle: str):
        try:
            return handle, ig.get("/instagram/profile", handle=handle)["data"]
        except LookupError:
            return handle, None      # private or missing — a result, not a crash

    with ThreadPoolExecutor(max_workers=workers) as pool:
        return dict(pool.map(one, handles))

requests.Session is not thread-safe for every use, but concurrent GETs with a shared connection pool are fine in practice. If you are running tens of thousands of handles, move to httpx.AsyncClient and drop the thread pool entirely.

Controlling cost

Three levers, in order of impact:

Control refresh frequency. Profile counts do not need the same cadence as comments on a new post. Pick intervals by how fast the data moves.

Only refresh aggressively when a human is waiting. Background jobs can run on a slower schedule; user-triggered refreshes can call the endpoint immediately:

ig.get("/instagram/profile", handle=handle)

Cap pagination. Every page is a credit. max_pages is not optional in production code.

Persisting a time series

The API returns snapshots. History is yours to keep, and it is where the actual value accumulates:

import sqlite3

db = sqlite3.connect("instagram.db")
db.execute("""
  create table if not exists snapshots (
    handle text, follower_count int, media_count int, captured_at text
  )
""")

for handle, profile in fetch_many(tracked_handles).items():
    if profile:
        db.execute(
            "insert into snapshots values (?, ?, ?, datetime('now'))",
            (handle, profile["follower_count"], profile["media_count"]),
        )
db.commit()

One credit per handle per day buys you a growth curve that no provider can take away from you later.

Common mistakes

Mistake What happens
Retrying a private-account lookup result Billed at the endpoint rate; the account is still private tomorrow.
Unbounded pagination The largest surprise credit charges we see, without exception.
Ignoring requested_at You cannot tell when a stored number was collected.
Refreshing everything too often Paying for data your product does not actually use.
Hardcoding the key It ends up in git. Use the environment.

Next: the runnable Python reference for every endpoint, the same thing in Node.js, or the endpoint documentation.

Try it with 100 free credits.

No credit card, credits never expire, and failed requests are not charged.