Blog / Comparisons
How to Scrape Instagram Data: Four Approaches
By the InScrape API team · Published 2026-09-02 · 7 min read

There are four ways to get Instagram data into a database, and the reason the advice online contradicts itself is that each of the four is genuinely correct for a different job. A dissertation that needs 800 posts once is not the same problem as a product that refreshes 5,000 profiles every morning, and the approach that wins the first loses the second badly.
Two variables decide it, in this order. Whose accounts? Accounts whose owners will log into your app, versus any public account on the platform — this one is binary and it eliminates an entire option immediately. Once, or forever? A single extraction has no maintenance cost, because nothing has to keep working next month. A pipeline has almost nothing but maintenance cost.
Answer those two and the rest of this article is mostly confirmation.
Option 1: Meta's own APIs
The Instagram Graph API and the Instagram API with Instagram Login are publishing and insights APIs. They read accounts that have completed an OAuth flow with your app, and nothing else.
That constraint is not a rate limit you can raise or a permission you can apply for. It is the design. If your use case is competitor tracking, creator discovery, hashtag monitoring or any form of market research, no amount of paperwork gets you there, because the accounts in question will never connect to your app.
Where it is not just correct but the only answer: reach, impressions, saves, profile views, audience demographics. These are private metrics. They exist only for the account owner, they come only from Meta, and any third party offering them for accounts they do not manage is estimating without saying so. We cannot supply them either, and it is worth being blunt about that up front.
Time to first call is days to weeks, and most of it is app review rather than engineering. The full side-by-side of what each surface returns is in Instagram Graph API vs scraping API.
Option 2: Browser automation with Playwright or Puppeteer
This is where most people start, because it works on the first afternoon. Launch a headless browser, load a public profile, read what a logged-out visitor sees. Then it degrades, in a predictable order.
The logged-out view narrows. A public profile is served to an anonymous visitor, then interrupted by a login interstitial after a handful of loads from the same address. The threshold is not published and it moves.
Selectors rot. Class names in the rendered DOM are hashed build artefacts, not stable identifiers. A deploy on Meta's side renames them and every CSS selector returns null, with no error and no warning — your pipeline succeeds and writes empty rows, which is worse than failing.
The payload you want is JSON, not DOM. It sits embedded in the page or behind a GraphQL request the page makes, keyed by a versioned document identifier that changes without notice. Pinning one is cleaner than DOM parsing right up until the day it 404s.
Datacentre IPs get challenged, and browsers are expensive. Cloud ranges are recognised, so volume means residential proxies. Each concurrent page is a full browser holding several hundred megabytes and pulling one to three megabytes per load, for fields an HTTP request would have returned as JSON.
None of this makes browser automation wrong. For a one-off extraction of a few thousand pages, run it from a residential connection over a weekend and you are done — no proxy contract, no maintenance, no bill. It is continuous operation that turns it into an infrastructure project you did not intend to start.
Option 3: Open-source libraries
Free, installable in a minute, and split into two families that carry very different risks.
Anonymous scrapers read the public web view without credentials. They are the safer family, and they inherit exactly the fragility described above — plus a rate limit low enough that most people abandon them somewhere in the first thousand requests.
Private API reimplementations replay the mobile app's internal endpoints. They return considerably more, and they need an account to log in with. An automated login gets challenged, and a checkpoint means a human solves a puzzle before your pipeline resumes. It also moves you across a line: logged-out reading of public pages is a much easier position to defend than automation inside an authenticated session, which is the whole subject of is Instagram scraping legal. The name is also routinely misread as access to private accounts, which it is not — see private Instagram API.
Either way, run three checks before building on one: date of the last commit, ratio of open to closed issues, and whether the recent issues are all some variant of "not working". A library that broke three months ago is not free. You have agreed to become its maintainer. Instagram scrapers on GitHub breaks the published projects into categories and says how each one fails.
Option 4: A hosted API
You pay money and someone else absorbs the breakage. The proxy pool, the parser updates, the retry logic and the 3am page still exist; they are just on the other side of an HTTP boundary.
curl "https://api.socialscrape.dev/v1/instagram/profile?handle=natgeo" \
-H "x-api-key: $INSCRAPE_KEY"
{
"success": true,
"credits_charged": 1,
"credits_remaining": 99,
"processing_time_ms": 1842,
"requested_at": "2026-09-13T14:32:18Z",
"query": { "username": "natgeo" },
"data": {
"handle": "natgeo",
"follower_count": 279000000,
"is_verified": true,
"media_count": 30412
}
}
No Instagram account is involved on either side. Every list endpoint returns next_cursor, which you pass back as cursor. Billing depends on actual scraper resource use, not the HTTP status. A private or missing target is charged if the scraper had to use an Instagram or H API resource to discover that result.
The honest downside is that you are now dependent on a vendor's uptime and pricing, for data you could fetch yourself. That trade is worth it only if your engineering time is worth more than the invoice, which is the arithmetic below.
The cost comparison, in hours
Dollar comparisons flatter self-hosting because the largest input is unpriced. These are our own figures from running this infrastructure; substitute your own where you have them.
| Meta Graph API | Browser automation | Open-source library | Hosted API | |
|---|---|---|---|---|
| Hours to first row | 8–40 | 4–8 | 1–2 | under 1 |
| Hours to something you would put in front of a customer | 20–40 | 80–160 | 40–80 | 2–4 |
| Maintenance, hours per month | 1–2 | 8–20 | 4–12 | 0 |
| Unplanned breakages per year | rare, and announced in advance | 4–12 | whenever upstream breaks and nobody merges a fix | not yours |
| Infrastructure to operate | none | browsers, queue, residential proxies | proxies, and accounts if it logs in | none |
| Who gets paged at 3am | nobody | you | you | your provider |
Price a concrete workload: 5,000 public profiles refreshed daily, roughly 150,000 requests per month.
Browser automation at 2 MB per load is around 300 GB of proxied traffic a month. Residential bandwidth sells by the gigabyte and the rate moves, so check today's and multiply: 300 GB is a monthly proxy line item on its own, before a single CPU, before a single hour of anyone's attention, and before the day and a half per month of steady-state engineering.
The same 150,000 requests at one credit each use 30% of the $299 Scale pack, which carries 500,000 credits and never expires. Be conservative when you budget: plan for every scheduled successful request to be billed, then use credits_charged in production logs to measure the real cost. The same arithmetic for every pricing model is in Instagram API cost.
The break-even is not close for a pipeline that runs continuously. It reverses entirely for a single extraction, where maintenance is zero by definition and a weekend of Playwright costs nothing but the weekend.
What no approach on this list can do
Worth stating plainly, because several of these get advertised:
- Private accounts. Not readable by any method here. A confirmed private profile returns 200 with available profile details, billed at the endpoint rate.
- Reach, impressions, saves, profile views. Owner-only, available exclusively through Meta's own APIs for connected accounts.
- Direct messages, notifications, saved posts, story viewer lists. No.
- Email addresses and phone numbers of followers. Not returned. A public business contact button is a different thing from a harvested contact list, and we do not supply the latter.
- Anything that writes. This is read-only. No posting, commenting, liking or following.
Derived metrics such as engagement rate, posting cadence or follower-quality signals are things you compute from the fields that come back. Any provider selling them is running the same arithmetic on the same public numbers.
Choosing
- Reporting on accounts your users connected → Meta's APIs, and nothing else will do
- One extraction, a few thousand pages, no ongoing obligation → Playwright from a residential connection
- Prototype or coursework, low volume, tolerance for breakage → an open-source library, after you check its commit history
- Anything a customer depends on, running daily → buy it, or accept that you have hired yourself as a scraping-infrastructure engineer
Where to go next
- The Instagram Data API overview, for the endpoint list and what each returns
- The documentation, for authentication, caching, pagination and the six error codes
- Pricing, to check the arithmetic above against your own volume
- Is Instagram scraping legal, before you choose anything that logs in
The free tier is 100 credits with no card, which is enough to run your real handles through and see what comes back before any of this becomes a decision.
Try it with 100 free credits.
No credit card, credits never expire, and failed requests are not charged.

