Is It Legal to Scrape Reddit and HN for Sales Leads?
Reddit's 2024 policy requires a contract for commercial data use and just sued Perplexity over it. Here's what CFAA and GDPR actually let you do.

You built the monitor, wired it to Slack, and for two weeks it's been quietly finding people on Reddit who are asking the exact question your product answers. Then a friend who happens to practice law asks, over drinks, how exactly you're p
You built the monitor, wired it to Slack, and for two weeks it's been quietly finding people on Reddit who are asking the exact question your product answers. Then a friend who happens to practice law asks, over drinks, how exactly you're pulling all this off Reddit. "The API, mostly," you say. She asks if you've actually read what Reddit's terms say about that. You haven't. Almost nobody running a Reddit or Hacker News lead pipeline has, because the honest answer to whether it's legal to scrape Reddit for sales leads spans four bodies of law that don't fully agree with each other, and guessing wrong costs more than most founders assume.
Reading a post and building a list from it are two different legal questions
"Is this legal" is really three separate questions wearing one trench coat. Question one: does reading a public post without logging in violate a computer-crime statute. Question two: does collecting, storing, and reusing that content commercially violate a contract you agreed to (or never agreed to) with the platform. Question three: does turning a username into a named, contactable person and emailing them trigger a privacy law. Each question has a different answer, a different court or regulator behind it, and a different level of risk. Conflating them is why founders either panic and shut the whole channel down, or don't panic at all and build on a foundation nobody checked.
Reddit drew a hard line in May 2024, and it's suing over it now
Reddit's Public Content Policy, reported in detail by TechCrunch when it went public on May 9, 2024, states that any business wanting to use Reddit data to "power, augment or enhance" a product for commercial purposes needs a contract. The same policy bars partners from using Reddit content to identify individuals for ad targeting, to spam or harass users, or to run background checks, facial recognition, or law-enforcement surveillance. Reddit has said the policy mostly formalized rules it already enforced internally, not a brand-new restriction sprung on the ecosystem overnight.
That policy isn't theoretical. On October 22, 2025, Reddit filed suit against Perplexity, Oxylabs, SerpApi, and AWMProxy in the Southern District of New York, covered by the National Law Review, alleging the group evaded Reddit's technical access controls at industrial scale to harvest content through Google's search index. The complaint leans on the anti-circumvention provisions of the DMCA rather than a garden-variety copyright claim, which matters because that theory targets bypassing an access control regardless of whether the underlying use might otherwise be defensible.
The federal scraping fight already has an answer, mostly
The question of whether scraping a public website is a computer crime got settled by hiQ Labs v. LinkedIn. The Ninth Circuit ruled, most recently in April 2022, that accessing data a site makes publicly available, no login, no bypassed wall, doesn't violate the Computer Fraud and Abuse Act, because there was no unauthorized access to begin with. That's a real, durable answer, and it's why nobody credible argues that reading a public Reddit thread is a federal crime.
A platform's terms of service still binds you even after you win the CFAA argument
This is where Reddit's Public Content Policy and the hiQ contract claim rhyme. Winning the CFAA argument tells you the government won't prosecute you for reading a public page. It says nothing about whether you agreed, by using the site, not to reuse that content the way you're using it. The Reddit lead generation playbook already draws the practical version of this line: reading public posts, classifying them, and notifying yourself is the normal, low-risk use most monitoring tools and DIY founders actually run. Naked promotional replies, sockpuppet accounts, and brigading are the behaviors that get accounts banned and, per Reddit's own policy language, edge toward exactly what the contract was written to stop. The same logic is why replying without getting banned is as much a legal hygiene habit as a growth tactic.
Winning the CFAA argument means nobody can call the police on you for reading a public post. It has never meant the terms of service stopped applying.
Hacker News never built the wall Reddit built
Hacker News runs the opposite playbook. Its Firebase and Algolia APIs are public, require no registration, and as of this writing carry no published rate limit and no commercial-use contract, a stark contrast to Reddit's Public Content Policy. That's one entire layer of legal risk that simply isn't there for HN monitoring the way it is for Reddit. The HN monitoring setup most founders run is built entirely on that open API for exactly this reason: it's the lower-friction channel, legally and technically. It doesn't mean anything goes. HN's moderators enforce the same social norms against sockpuppeting and naked self-promotion that Reddit's communities do; there's just no separate commercial-data contract sitting on top of it.
GDPR doesn't care that the post was public
Public doesn't mean unregulated. GDPR applies to the personal data of anyone in the EU regardless of whether you found it on a public forum, and it kicks in the moment you store an identifiable person's details for reuse, not when you merely read their post. The ICO's own guidance is direct on this: direct marketing can rely on legitimate interest as a lawful basis, but "legitimate interests can't legitimise unlawful processing," and for electronic channels like email, PECR-style consent rules sit on top of that basis and can't be overridden by it.
The practical line for a founder running this today
Strip out the jurisdiction-hopping and one line survives across all four bodies of law: reading, at human or API-rate-limited scale, for your own notification, is the low-risk end of this spectrum. Building a product that ingests and resells Reddit's data at scale is what triggers the commercial-contract requirement, and Reddit has now shown twice, once as policy language and once as an actual federal complaint, that it will enforce that line. GDPR shows up not when you read a post but when you store a person's identity for outreach, and the follow-up channel, email especially, carries its own consent rules on top. None of the four laws say the same thing, which is exactly why "is Reddit scraping legal" gets a different answer depending on which part of the pipeline you're asking about.
● FAQ
- Is it illegal to read public Reddit posts to find sales leads?
- No. The Ninth Circuit ruled in hiQ v. LinkedIn that scraping data a website makes publicly available, without logging in or bypassing a wall, doesn't violate the Computer Fraud and Abuse Act, because there was no unauthorized access to begin with. That answers the federal criminal-access question. It doesn't answer whether you've broken a contract with the platform, which is a separate question with a separate answer.
- Do I need a commercial agreement with Reddit to monitor it for leads?
- Depends what you're building. Reddit's Public Content Policy requires a contract for anyone using Reddit data to power, augment, or enhance a product commercially, which is squarely aimed at companies reselling or training on Reddit's data at scale. A founder reading threads through the standard developer API, under their own rate-limited account, to notify themselves of relevant posts, is a much smaller footprint than what that policy was written to catch. It's a real distinction, but it's not a bright line, and Reddit has shown in 2025 it will litigate the ambiguous cases.
- Does GDPR apply if I'm just reading a post, not collecting anyone's data?
- Reading doesn't trigger GDPR. Storing a username, an inferred email, or any other identifier tied to a real EU person, in a list you'll use for outreach, does. At that point you need a lawful basis, legitimate interest is the usual one for B2B prospecting, and the person has the right to object and have you stop. Following up by email adds a second, stricter layer: PECR-style consent rules for direct electronic marketing, which legitimate interest alone doesn't satisfy.
- Is Hacker News less risky than Reddit for this?
- Structurally, yes. HN's Firebase and Algolia APIs are free, require no registration, and carry no published rate limit or commercial-use contract, unlike Reddit's Public Content Policy. That removes one entire layer of risk. It doesn't remove GDPR if your leads include EU citizens, and it doesn't remove HN's own social norms against sockpuppeting or naked self-promotion, which are enforced by moderators, not lawyers.
Three more from the log.

How to reply on Reddit without getting banned
Reddit reply strategy for founders: why most marketing advice gets you banned, how moderators actually think, and the disclosure pattern that earns upvotes.
Jan 09, 2026 · 10 min
The Buying-Intent Window Is 16 Minutes
Original data from 975 Ask HN posts and 20 subreddits: median time to first reply on a buying-intent post was 16 minutes, 92 percent were answered inside an hour, and every intent post on Reddit came from an operator community, none from a builder one.
Aug 25, 2026 · 4 min
How to find clients on Reddit (without getting banned)
Find paying clients on Reddit by being early, useful, and specific — not by spamming DMs. The honest 2026 playbook for freelancers and agencies.
May 04, 2026 · 5 min