Blog / Risk and compliance
Is Instagram Scraping Legal? Law vs Terms
By the InScrape API team · Published 2026-08-18 · 5 min read

"Is scraping legal" is three questions wearing one coat, and the answers point in different directions. That is why every thread on this ends in an argument: two people are answering two different questions and both are right.
Split them and it becomes tractable.
- Is it a crime? — computer misuse law, in the US the Computer Fraud and Abuse Act
- Is it a breach of contract? — Instagram's terms of service
- Is the data lawful for you to hold and use? — data protection law, GDPR and its relatives
Nothing below is legal advice. It is the map of where the questions live, so you can ask your own lawyer something specific instead of something vague.
Question one: is it a crime?
This is the question people are actually frightened of, and it is the one with the clearest answer for public data.
The CFAA criminalises accessing a computer "without authorization" or in a way that "exceeds authorized access". For years, companies argued that a cease-and-desist letter or a line in the terms of service was enough to turn scraping into unauthorised access.
Two decisions pushed back hard.
hiQ Labs v. LinkedIn ran through the Ninth Circuit and landed on a simple proposition: data that is publicly available, with no login required, is not accessed "without authorization" in the CFAA sense. There is no gate to break when the door is open to the whole internet.
Van Buren v. United States was the Supreme Court reading "exceeds authorized access" narrowly — a gates-up-or-down question. Do you have permission to be in this area of the system at all? Using data you were allowed to reach for a purpose someone dislikes is not the same as breaking in.
The practical upshot for public Instagram data: reading a profile page that any logged-out visitor can load is a poor candidate for a computer-crime theory. That is a very different statement from "scraping is legal", and the gap between the two is where people get hurt.
Question two: are you breaching a contract?
Here the picture changes, and it changes on one specific fact.
The terms of service are a contract. A contract binds people who agreed to it. If you created an account, clicked through the terms, and then ran automation against the platform while logged in, you agreed to something and you are plausibly breaching it.
If you never had an account and never agreed to anything, the argument that you are bound is much weaker. Meta Platforms v. Bright Data turned on close to this point: the court was unpersuaded that logged-out scraping breached the user agreement, because the agreement governs what a user does, and the scraping in question happened without being logged in.
So the line that matters is not "scraping vs not scraping". It is:
| Logged out, public pages | Logged in, or using an account's session | |
|---|---|---|
| Contract exposure | Weak — you agreed to nothing | Real — you accepted the terms |
| CFAA exposure | Weak — no access gate | Stronger — you passed an authentication gate |
| Platform enforcement | Rate limits, blocks | Account restriction, disablement, legal notice |
This is also, incidentally, the reason a scraping library that asks for your username and password is a worse deal than it looks. It is not only an operational risk. It moves you across the line on both of the first two questions at once.
Question three: is the data lawful to hold?
This is the one that gets skipped, and it is the one most likely to actually cost a company money.
Public does not mean unregulated. Under GDPR, a name, a handle, a photograph of an identifiable person and a bio are personal data whether or not the person published them openly. There is no general "it was publicly available" exemption. You still need a lawful basis — for commercial scraping that is usually legitimate interests, which requires you to actually run and document the balancing test, not assert it.
There is also Article 14, the notice obligation for data you collected from somewhere other than the person themselves. Its practical difficulty is the point: telling several million people you hold data about them is hard, and the exemptions are narrower than most teams assume.
CCPA has a "publicly available" carve-out, but it is drafted around information lawfully made available from government records and similar sources. It is not a blanket exemption for anything visible on a website.
None of this makes the data untouchable. It makes it governed. The teams that get in trouble are the ones who treated question three as though answering questions one and two had covered it.
Where it actually goes wrong
In practice, the trouble almost never comes from "we read public profile data". It comes from four specific patterns:
Scraping while authenticated. You handed a tool your credentials, or you built on a library that logs in. You have now agreed to the terms and passed an access gate. Both of the first two questions get harder.
Touching private accounts. A private account is an access control. Circumventing it is exactly the "gate" that Van Buren left intact. There is no reading of the case law that helps you here.
Harvesting contact details. Email and phone harvesting attracts a different and much less forgiving body of law — anti-spam regimes, and data protection rules that treat contact data as directly identifying. "It was on the profile" is not the defence people think it is.
Building a product that re-identifies or profiles individuals. Aggregating public posts into a dossier about a private person is a different activity from counting engagement on brand accounts, and regulators read it differently.
Notice that none of these are about the mechanics of fetching a page. They are all about what you fetch and what you do with it.
The version to build
If you want the shortest defensible position:
- Read only what a logged-out visitor can see
- Never authenticate, never use anyone's session or credentials
- Never attempt private accounts — treat a private-account lookup result as final
- Skip contact-detail fields even when they are exposed
- Write down your lawful basis and your legitimate-interests balancing test before you launch, not after a complaint
- Retain for a defined period, and honour deletion requests
- Keep the aggregate, not the dossier
That is not a legal opinion, it is a risk posture — but it is the posture that keeps the first two questions easy and leaves you with only the third to manage properly.
The Instagram Data API is built to that posture on purpose: no login, no credentials, no session, and a 200 private-account lookup result charged at the endpoint rate. It does not answer question three for you — nothing can, because that answer depends on what you build — but it means questions one and two are not the ones keeping you up.
Try it with 100 free credits.
No credit card, credits never expire, and failed requests are not charged.

