Blog / Comparisons
Instagram Scrapers on GitHub: What Breaks
By the InScrape API team · Published 2026-09-02 · 8 min read

Search GitHub for an Instagram scraper and you get several hundred repositories that all claim to work. Most of them do work, on the day their author last ran them, from a home IP, against a profile nobody was watching. The useful question is not which repo is best. It is which category you are picking up, because the failure mode is a property of the category rather than the repo.
There are four. No star counts or commit dates below, deliberately: they age badly and they mislead. A popular repo with no commits since Instagram's last app change is worse than an obscure one merged last week. This page is only about the open-source category; the four ways to get Instagram data at all is the wider decision, and picking up a repo is one branch of it.
Category 1: media downloaders
The Instaloader shape. A command-line tool that walks a public profile and writes posts, captions, comments and story files to disk, usually with a JSON sidecar per item and the ability to resume where it stopped.
Good at: archiving. For "keep a copy of this account's posts" nothing else comes close, and resumption after an interruption is handled properly, which hand-rolled scripts never get around to.
Where it breaks: it is a downloader, not a data source. You get a directory tree, then you write and maintain the code that turns files into rows. Anonymous mode is throttled hard and hits login walls beyond shallow paging, so the docs nudge you toward --login — which puts you in category 2 with all of its consequences.
Category 2: private-API wrappers
The instagrapi shape, and the older instagram-private-api ports. These reimplement the endpoints Instagram's mobile app calls, including the device fingerprinting and request signing that make the traffic look like a phone.
Good at: breadth. Speaking the app's protocol reaches things the logged-out web never exposes, and these libraries can act as well as read — follow, like, post, message. For automating an account you own, this category is the usual open-source answer, a logged-in headless browser being the other.
Where it breaks: all of it requires credentials, and that is the whole story.
Logging into a real Instagram account in order to scrape is the fastest way to lose that account. Not a risk to manage, the expected outcome. The sequence is consistent even where the volume that triggers it is not: a challenge_required response, a checkpoint wanting a code from the email or phone on file, then either a human solves it or the pipeline is down. Repeat that from a datacentre IP and the account is disabled. Burners do not help — fresh accounts with no device history are challenged sooner than established ones, and buying aged accounts adds fraud to your problem list rather than removing anything.
These libraries also break most often, because they track a moving internal protocol — what a private Instagram API actually is walks through the session, fingerprint and signing work that keeps breaking. Read the issue tracker before depending on one: if the newest open issues are variants of "login broken" and the last release predates them, the repo is abandoned and not yet labelled as such. Authenticating changes your legal position too, not just your uptime — logged-out reading and session-based reading are different activities under both the terms of service and computer-misuse law, which is the subject of is Instagram scraping legal.
Category 3: headless-browser scrapers
Playwright, Puppeteer or Selenium driving a browser at instagram.com, reading the rendered DOM or intercepting the XHR responses the page makes.
Good at: fidelity and debugging. Whatever a logged-out visitor sees, you see, and you can watch it fail in a headed browser. Intercepting the network responses instead of the DOM gives you the same JSON the page consumes, which survives markup changes.
Where it breaks: economics first, detection second. A browser per request is one to three seconds and a few hundred megabytes of RAM, where the same fields arrive from a single HTTP request holding no process at all, so concurrency becomes a fleet. Datacentre ranges get a login wall after a handful of loads, so you add residential proxies, then session reuse to amortise them, then a queue, then health checks on the queue — and the feature you meant to build is still untouched.
Category 4: thin HTTP scrapers
A requests call to one of the public web endpoints — the ?__a=1-era profile JSON, the web profile info route, or a GraphQL call with a query_hash copied out of the network tab — and a dictionary walk over the response.
Good at: being fifty lines. No browser, no login, no dependencies past an HTTP client. For one public profile a few times a day this is the correct amount of code, and everyone should stop there.
Where it breaks: it is built on constants that are not yours. Query hashes rotate, response shapes change without notice, and routes that served anonymous callers for years get gated behind a session with no announcement. Rate limiting is per-IP and arrives fast, so a loop over a handle list gets a 401 or an empty payload partway through and, unchecked, writes nulls into your database.
The four in one table
| Needs an account | Needs proxies | Latency per item | What ends it | |
|---|---|---|---|---|
| Media downloader | For anything deep | Past small volumes | Seconds | It is a CLI, and login walls |
| Private-API wrapper | Always | Yes | Fast while it works | Challenges, then a ban |
| Headless browser | No, for public pages | Yes | 1–3s | Cost, then anti-bot |
| Thin HTTP client | No | Immediately | One request | Rotated hashes, IP limits |
Which data is hardest to keep working
No public rate-limit table exists for these endpoints. Anyone quoting exact numbers is reporting a measurement of their own IP on one particular day. The ordering, though, is stable across all four approaches:
- A single public profile — one request, the last thing to break
- Posts past the first page — the cursor is where anonymous access starts getting refused
- Search and hashtag listings — throttled hardest, being the cheapest to abuse
- Followers and following — deep cursor paging, which is why the open-source libraries that offer it are almost all category 2 libraries
- Stories and highlights — ephemeral and session-gated
If you only need item 1, a library is very likely the right answer. If you need 3 to 5, you are being pushed toward credentials, and it is worth noticing that before writing the code rather than after.
Vetting a repo before you depend on it
- Last commit measured against Instagram's app releases, not against today
- The newest ten open issues, not the README
- A hardcoded
query_hashor app version in source is a fixed constant against a moving target - No documented proxy support means it has never run at volume
When a free library is the right call
Genuinely often: one-off research you can rerun by hand, under a few hundred profiles a month with no uptime requirement, archiving an account you own, or learning how any of this works — which needs no further justification. A script that costs nothing and breaks twice a year is the correct answer when a human is standing next to it.
When it stops making sense
The line is not volume, it is who notices the failure. Once a customer sees the output, maintenance becomes support, and the arithmetic moves fast: one engineer-hour a month on rotated hashes and challenge screens already costs more than the Starter plan. Add proxies, which every category above needs past trivial volumes, and a self-hosted scraper is not free, it is unbilled.
The other signal is the shape of the data. Cursor paging past the first page, concurrency across thousands of handles, a response contract stable enough to write a schema against — at that point you are not using a library, you are staffing one. That is all a hosted API sells: the proxy pool, the parser maintenance and the retries exist either way, they are just on somebody else's side of the boundary.
curl "https://api.socialscrape.dev/v1/instagram/profile?handle=natgeo" \
-H "x-api-key: $INSCRAPE_KEY"
Every successful endpoint answers with the same envelope — success, credits_charged, credits_remaining, processing_time_ms, requested_at, query, data — and confirmed private accounts return 200 with available profile details, charged at the endpoint rate. Most endpoints are one credit; followers and following are two.
What we deliberately do not do
Worth stating, because half the repositories above will attempt things we cannot and will not:
- No login, ever. No credentials, no session, no account to lose. Which also means no direct messages, no notifications, no saved posts, and no insights, reach or impressions.
- No writes. No posting, commenting, liking or following. The surface is read-only, so nothing built on it can act as an account.
- No private accounts. Confirmed private accounts return 200 with available profile details, billed at the endpoint rate. No flag changes that.
- No contact harvesting. Follower lists come back as handles and public profile fields, not emails or phone numbers.
- Derived metrics are yours. Engagement rate, posting cadence, suspicious-follower ratios: we return the counts, you do the arithmetic.
If your requirement is on that list, a category 2 library with an account you are willing to lose is the only thing that will do it. Read what triggers Instagram's scraping warning before you spend it.
Where to go next
- Instagram API without login — the same options judged on the credential question alone
- Instagram API with Python — pagination, retries and cost control once you have picked an approach
- The followers endpoint — the one that pushes every open-source route toward a login
Try it with 100 free credits.
No credit card, credits never expire, and failed requests are not charged.

