TL;DR
7 min readScraping publicly visible Reddit pages is not a federal crime in the United States after Van Buren and hiQ, but it breaches the Reddit User Agreement, which prohibits scraping the Services without Reddit's prior written consent and grants crawl permission only within robots.txt, where reddit.com currently sends Disallow slash to every user agent. Reddit enforces that contract with token revocation, account and app suspension, audit rights and data licensing deals. The compliant route is the official Data API after approval, with a separate written agreement for any commercial use.
Is scraping Reddit legal?
No United States statute answers "is scraping Reddit legal", so Reddit answers it by contract: the Reddit User Agreement prohibits scraping the Services without Reddit's prior written consent and grants permission to crawl only within the parameters of its robots.txt file. Criminal computer access law and contract law decide separately, and Reddit enforces on the contract side, where it does not need a prosecutor. Read Reddit's own terms for your use case before you write a client, because those terms bind your account whatever the criminal answer is.
The general law here, including what Van Buren v. United States held about "exceeds authorized access" and what hiQ Labs v. LinkedIn held about scraping public pages under the Computer Fraud and Abuse Act, sits on the page about when web scraping is legal. This page covers what Reddit, Inc. requires. Neither is legal advice.
What does the Reddit User Agreement say about scraping?
The Reddit User Agreement bans scraping outright and permits crawling only conditionally. Section 7, "Things You Cannot Do", says you may not "Access, search, or collect data from the Services by any means (automated or otherwise) except as permitted in these Terms or in a separate agreement with Reddit (we conditionally grant permission to crawl the Services in accordance with the parameters set forth in our robots.txt file, but scraping the Services without Reddit's prior written consent is prohibited)." That text comes from the User Agreement effective July 1, 2026, last revised May 26, 2026, at redditinc.com/policies/user-agreement.
The prohibition covers collection "by any means (automated or otherwise)", so a headless browser and a curl loop land in the same clause. Reddit lists section 7 among the sections that survive termination, so deleting the account you scraped from does not retire the obligation.

Does Reddit robots.txt permit crawling?
Reddit's robots.txt currently disallows every crawler from every path. The file at reddit.com/robots.txt contains one rule group, "User-agent: *" followed by "Disallow: /", under a header comment that reads "Reddit believes in an open internet, but not the misuse of public content" and points to Reddit's Public Content Policy. The conditional crawl permission in section 7 therefore resolves to no permission for a general crawler today.
A robots.txt file has no force of its own, as the robots.txt explainer sets out. Reddit's User Agreement converts it into a contractual boundary by granting crawl permission by reference to whatever the file says when you read it.
What counts as commercial use of Reddit data?
Reddit defines commercial use as any use by a business, on behalf of a business, or as part of a monetized product or service. The help article "Developer Platform and Accessing Reddit Data" lists mobile apps with ads or paywalls, subscription services, services or data access for fees, publishing Reddit content on monetized sites, and selling access to models trained on Reddit data. The same article states you cannot display Reddit content and run advertisements in your app, website or service.
Both contracts route commercial use to a signed deal. Data API Terms 3.1 says that for "commercial purposes, research in excess of rate limits, or for any use that is not expressly permitted", you "will need to enter into a separate agreement with Reddit". Developer Terms 4.1 repeats it and bars access "by or on behalf of a business or as part of a service or product that is monetized" without written approval.
| Route to Reddit data | What Reddit's policy says | Primary source |
|---|---|---|
| Scraping public pages | Prohibited without Reddit's prior written consent | User Agreement, section 7 |
| Crawling under robots.txt | Conditionally permitted, and the live file sends Disallow: / to every user agent | User Agreement section 7, reddit.com/robots.txt |
| Data API, non-commercial | Allowed after you request access and Reddit approves the use case | Responsible Builder Policy |
| Data API, commercial | Requires a separate written agreement with Reddit | Data API Terms 3.1, Developer Terms 4.1 |
| Academic research | Only through the Reddit for Researchers program | Developer Platform and Accessing Reddit Data |
| Training an AI model | Prohibited without Reddit's permission and rightsholder permission | Data API Terms 2.4, Developer Terms 4.2 |
What changed for Reddit data access in 2023?
Reddit issued new Data API Terms effective June 19, 2023, and that document is still in force, last revised July 20, 2026. It gives Reddit the right to "set and enforce limits on your use of the Data APIs" in its sole discretion under section 2.9, bars you from circumventing those limits under section 3.2, and requires deletion of cached User Content on termination under section 6. The breakdown of Reddit API pricing covers the rates that followed, and Reddit API rate limits covers the quota mechanics your client has to respect.
Pushshift, the archive most Reddit research ran on, stopped serving the general public in the same period. Its dump series ends at March 2023, which is where the Arctic Shift archive picks it up. Pushshift now gates access behind a moderator check: its signup terms require the user to certify that they are a registered Reddit moderator and limit access to "community moderation, enforcing Reddit community guidelines, and ensuring community member safety".
How does Reddit enforce its terms?
Reddit enforces through access controls, audits and licensing rather than through criminal referral. The Responsible Builder Policy lists the enforcement actions plainly: revoking your access tokens, suspending your app or account, and suspending associated accounts, bots, domains or subreddits. Developer Terms 2.2 adds an audit right: Reddit monitors your use, requires evidence of compliance, and may "use any technical measures to prevent, block, or otherwise overcome the interference".
The licensing side comes from the Public Content Policy. Reddit says it enters licensing arrangements so it can know who is accessing public content and why, place contractual restrictions on use, and ensure licensees honor deletions by Redditors. That policy names unauthorized access "by scraping or using data brokers" as the behavior the licensing program exists to displace.

What is the compliant way to collect Reddit data?
Register an app, request Data API access for a stated use case, and sign a separate agreement with Reddit before any money touches the output. The Responsible Builder Policy sets three conditions on all Reddit data access: approval before you access anything, no masking how and why you are accessing data, and no circumventing or exceeding access limits. Data API Terms 2.8 makes the second concrete, requiring the OAuth token Reddit issued and no masking of the user agent or OAuth identity.
RedReplier
Get Started
Reddit, X, Bluesky & HN
Real-time intent alerts
Unlimited AI replies
Ranked by buyer intent
Three obligations follow. Data API Terms 3.2 requires you to delete any data not required for the approved use case. Developer Terms 3.3 requires you to remove User Content once Reddit or its author deletes it. Data API Terms 2.4 grants no training right, so a model needs its own permission. If that engineering is not worth building, buy the access: RedReplier compared against building on the Reddit API sets out the tradeoff, and Reddit keyword monitoring delivers the same conversations as alerts.
Frequently Asked Questions
Is scraping Reddit legal if the pages are public?
Public visibility settles the criminal question and not the contract question. US courts stopped treating access to public pages as a Computer Fraud and Abuse Act violation, which the general web scraping page covers, but Reddit's User Agreement prohibits scraping without prior written consent whether or not a page requires a login. Reddit's Public Content Policy says the same: third parties have no right to misuse public content just because it is public.
Does Reddit robots.txt block all crawlers?
Yes. The live file at reddit.com/robots.txt contains a single group, "User-agent: *" with "Disallow: /", so it names no exception for any crawler. Reddit's User Agreement grants crawl permission only "in accordance with the parameters set forth in our robots.txt file", so the current file leaves that grant empty.
Can I use the Reddit Data API inside a paid product?
Not without a separate agreement. Data API Terms 3.1 routes commercial purposes to a separate agreement with Reddit, and Developer Terms 4.1 bars access by or on behalf of a business or as part of a monetized product without written approval. Reddit's help documentation counts subscription services, paywalled apps and data access for fees as commercial.
Can I train an AI model on Reddit posts?
Not without permission from Reddit and from the rightsholders. Data API Terms 2.4 grants no right to use User Content "for training a machine learning or AI model, without the express permission of rightsholders", and Developer Terms 4.2 bars accessing Reddit data by API, indexing, caching or crawling to train models without Reddit's permission. Reddit's help documentation adds that commercial use of a model trained on Reddit data needs explicit approval.
Is Pushshift still available?
Pushshift now serves Reddit moderators only. Its signup terms require the user to certify that they are a registered Reddit moderator and limit use to community moderation, guideline enforcement and member safety. Its public dump series ends at March 2023, and Arctic Shift continues the series from there.
What happens if Reddit catches you scraping?
Reddit revokes access tokens, suspends the app or account, and suspends associated accounts, bots, domains or subreddits, which the Responsible Builder Policy lists as its enforcement actions. Data API Terms 3.2 adds a permanent block on Data API access where Reddit believes you breached the limits clause. Because the breach is contractual, Reddit acts on its own authority and does not wait for a court.
What happens to my access if I close my Reddit account after scraping?
Closing the account does not retire the obligation. Reddit lists section 7, the clause banning scraping, among the sections that survive termination. The contract claim rests on the act of scraping, not on whether the account still exists.
How does Reddit decide who gets a data licensing deal?
Reddit's Public Content Policy frames licensing as a way to control access rather than a market open to anyone. Reddit enters these arrangements to know who is accessing public content and why, to place contractual restrictions on use, and to hold licensees to honoring deletions Redditors make. The policy names unauthorized scraping and data brokers as the behavior licensing is meant to replace.
Can university researchers get special Reddit data access?
Academic research runs through a separate channel from the standard Data API. Reddit routes it through the Reddit for Researchers program rather than general API approval or scraping. Scraping and unapproved API use stay off limits for research the same as for any other use case.
Can Reddit change Data API rate limits whenever it wants?
Reddit keeps that decision to itself. Data API Terms 2.9 gives Reddit the right to set and enforce limits on Data API use in its sole discretion, and section 3.2 bars developers from working around whatever limit is currently in force. A client built against today's limits can find itself throttled or blocked tomorrow with no negotiation.
See us more often in Google
One click marks RedReplier as a preferred source, so our articles sit higher in your Top Stories, AI Mode, and AI Overviews.
Before you go...
RedReplier
Catch every buyer asking for what you sell
RedReplier watches Reddit, X, Bluesky and Hacker News in real time, ranks every thread by buyer intent, and drafts your reply, so you get there first.
Reddit, X, Bluesky & HN
Real-time intent alerts
Unlimited AI replies
Ranked by buyer intent
Related Articles


When is web scraping legal? Four questions, four separate answers
No single rule answers "is web scraping legal": the CFAA, contract law, copyright and the GDPR each decide separately, and hiQ v. LinkedIn split two.


Inside Reddit API rate limits: 100 queries a minute, averaged over ten
How Reddit API rate limits work in practice: 100 queries per minute per OAuth client id, a rolling ten minute average, three X-Ratelimit headers, and 429.


Every Reddit Engagement Rate You See Was Built by Somebody Else
A Reddit engagement rate exists in no Reddit product. What the API returns, what post insights show, and the three rates people build out of those parts.

