glossary

When is scraping Reddit legal? The User Agreement answers before any court does

Taras Shynkarenko
Taras Shynkarenko
•Updated: •7 min read
When is scraping Reddit legal? The User Agreement answers before any court doesWhen is scraping Reddit legal? The User Agreement answers before any court does

TL;DR

7 min read

Scraping publicly visible Reddit pages is not a federal crime in the United States after Van Buren and hiQ, but it breaches the Reddit User Agreement, which prohibits scraping the Services without Reddit's prior written consent and grants crawl permission only within robots.txt, where reddit.com currently sends Disallow slash to every user agent. Reddit enforces that contract with token revocation, account and app suspension, audit rights and data licensing deals. The compliant route is the official Data API after approval, with a separate written agreement for any commercial use.

No United States statute answers "is scraping Reddit legal", so Reddit answers it by contract: the Reddit User Agreement prohibits scraping the Services without Reddit's prior written consent and grants permission to crawl only within the parameters of its robots.txt file. Criminal computer access law and contract law decide separately, and Reddit enforces on the contract side, where it does not need a prosecutor. Read Reddit's own terms for your use case before you write a client, because those terms bind your account whatever the criminal answer is.

The general law here, including what Van Buren v. United States held about "exceeds authorized access" and what hiQ Labs v. LinkedIn held about scraping public pages under the Computer Fraud and Abuse Act, sits on the page about when web scraping is legal. This page covers what Reddit, Inc. requires. Neither is legal advice.

What does the Reddit User Agreement say about scraping?

The Reddit User Agreement bans scraping outright and permits crawling only conditionally. Section 7, "Things You Cannot Do", says you may not "Access, search, or collect data from the Services by any means (automated or otherwise) except as permitted in these Terms or in a separate agreement with Reddit (we conditionally grant permission to crawl the Services in accordance with the parameters set forth in our robots.txt file, but scraping the Services without Reddit's prior written consent is prohibited)." That text comes from the User Agreement effective July 1, 2026, last revised May 26, 2026, at redditinc.com/policies/user-agreement.

The prohibition covers collection "by any means (automated or otherwise)", so a headless browser and a curl loop land in the same clause. Reddit lists section 7 among the sections that survive termination, so deleting the account you scraped from does not retire the obligation.

A developer's hands on a laptop keyboard, next to the section on whether Reddit's robots.txt permits crawling.

Does Reddit robots.txt permit crawling?

Reddit's robots.txt currently disallows every crawler from every path. The file at reddit.com/robots.txt contains one rule group, "User-agent: *" followed by "Disallow: /", under a header comment that reads "Reddit believes in an open internet, but not the misuse of public content" and points to Reddit's Public Content Policy. The conditional crawl permission in section 7 therefore resolves to no permission for a general crawler today.

A robots.txt file has no force of its own, as the robots.txt explainer sets out. Reddit's User Agreement converts it into a contractual boundary by granting crawl permission by reference to whatever the file says when you read it.

What counts as commercial use of Reddit data?

Reddit defines commercial use as any use by a business, on behalf of a business, or as part of a monetized product or service. The help article "Developer Platform and Accessing Reddit Data" lists mobile apps with ads or paywalls, subscription services, services or data access for fees, publishing Reddit content on monetized sites, and selling access to models trained on Reddit data. The same article states you cannot display Reddit content and run advertisements in your app, website or service.

Both contracts route commercial use to a signed deal. Data API Terms 3.1 says that for "commercial purposes, research in excess of rate limits, or for any use that is not expressly permitted", you "will need to enter into a separate agreement with Reddit". Developer Terms 4.1 repeats it and bars access "by or on behalf of a business or as part of a service or product that is monetized" without written approval.

Route to Reddit dataWhat Reddit's policy saysPrimary source
Scraping public pagesProhibited without Reddit's prior written consentUser Agreement, section 7
Crawling under robots.txtConditionally permitted, and the live file sends Disallow: / to every user agentUser Agreement section 7, reddit.com/robots.txt
Data API, non-commercialAllowed after you request access and Reddit approves the use caseResponsible Builder Policy
Data API, commercialRequires a separate written agreement with RedditData API Terms 3.1, Developer Terms 4.1
Academic researchOnly through the Reddit for Researchers programDeveloper Platform and Accessing Reddit Data
Training an AI modelProhibited without Reddit's permission and rightsholder permissionData API Terms 2.4, Developer Terms 4.2

What changed for Reddit data access in 2023?

Reddit issued new Data API Terms effective June 19, 2023, and that document is still in force, last revised July 20, 2026. It gives Reddit the right to "set and enforce limits on your use of the Data APIs" in its sole discretion under section 2.9, bars you from circumventing those limits under section 3.2, and requires deletion of cached User Content on termination under section 6. The breakdown of Reddit API pricing covers the rates that followed, and Reddit API rate limits covers the quota mechanics your client has to respect.

Pushshift, the archive most Reddit research ran on, stopped serving the general public in the same period. Its dump series ends at March 2023, which is where the Arctic Shift archive picks it up. Pushshift now gates access behind a moderator check: its signup terms require the user to certify that they are a registered Reddit moderator and limit access to "community moderation, enforcing Reddit community guidelines, and ensuring community member safety".

How does Reddit enforce its terms?

Reddit enforces through access controls, audits and licensing rather than through criminal referral. The Responsible Builder Policy lists the enforcement actions plainly: revoking your access tokens, suspending your app or account, and suspending associated accounts, bots, domains or subreddits. Developer Terms 2.2 adds an audit right: Reddit monitors your use, requires evidence of compliance, and may "use any technical measures to prevent, block, or otherwise overcome the interference".

The licensing side comes from the Public Content Policy. Reddit says it enters licensing arrangements so it can know who is accessing public content and why, place contractual restrictions on use, and ensure licensees honor deletions by Redditors. That policy names unauthorized access "by scraping or using data brokers" as the behavior the licensing program exists to displace.

A person reviewing a printed contract at a desk, next to the section on the compliant way to collect Reddit data.

Reddit's enforcement ladder
1
Revoke access tokens. Reddit pulls the OAuth tokens first, per the Responsible Builder Policy.
2
Suspend the app or account. Access for that specific app or account gets cut next.
3
Suspend associated accounts, bots, domains or subreddits. Connected accounts, bots, domains and subreddits go down too.
4
Block Data API access permanently. Data API Terms 3.2 adds a permanent block where Reddit believes you breached the limits clause.
Four steps Reddit takes to punish a breach of the scraping clause, from section 7 to a permanent block.

What is the compliant way to collect Reddit data?

Register an app, request Data API access for a stated use case, and sign a separate agreement with Reddit before any money touches the output. The Responsible Builder Policy sets three conditions on all Reddit data access: approval before you access anything, no masking how and why you are accessing data, and no circumventing or exceeding access limits. Data API Terms 2.8 makes the second concrete, requiring the OAuth token Reddit issued and no masking of the user agent or OAuth identity.

RedReplier
RedReplier

Get Started

Reddit, X, Bluesky & HN

Real-time intent alerts

Unlimited AI replies

Ranked by buyer intent

Three obligations follow. Data API Terms 3.2 requires you to delete any data not required for the approved use case. Developer Terms 3.3 requires you to remove User Content once Reddit or its author deletes it. Data API Terms 2.4 grants no training right, so a model needs its own permission. If that engineering is not worth building, buy the access: RedReplier compared against building on the Reddit API sets out the tradeoff, and Reddit keyword monitoring delivers the same conversations as alerts.

Frequently Asked Questions

Public visibility settles the criminal question and not the contract question. US courts stopped treating access to public pages as a Computer Fraud and Abuse Act violation, which the general web scraping page covers, but Reddit's User Agreement prohibits scraping without prior written consent whether or not a page requires a login. Reddit's Public Content Policy says the same: third parties have no right to misuse public content just because it is public.

Does Reddit robots.txt block all crawlers?

Yes. The live file at reddit.com/robots.txt contains a single group, "User-agent: *" with "Disallow: /", so it names no exception for any crawler. Reddit's User Agreement grants crawl permission only "in accordance with the parameters set forth in our robots.txt file", so the current file leaves that grant empty.

Can I use the Reddit Data API inside a paid product?

Not without a separate agreement. Data API Terms 3.1 routes commercial purposes to a separate agreement with Reddit, and Developer Terms 4.1 bars access by or on behalf of a business or as part of a monetized product without written approval. Reddit's help documentation counts subscription services, paywalled apps and data access for fees as commercial.

Can I train an AI model on Reddit posts?

Not without permission from Reddit and from the rightsholders. Data API Terms 2.4 grants no right to use User Content "for training a machine learning or AI model, without the express permission of rightsholders", and Developer Terms 4.2 bars accessing Reddit data by API, indexing, caching or crawling to train models without Reddit's permission. Reddit's help documentation adds that commercial use of a model trained on Reddit data needs explicit approval.

Is Pushshift still available?

Pushshift now serves Reddit moderators only. Its signup terms require the user to certify that they are a registered Reddit moderator and limit use to community moderation, guideline enforcement and member safety. Its public dump series ends at March 2023, and Arctic Shift continues the series from there.

What happens if Reddit catches you scraping?

Reddit revokes access tokens, suspends the app or account, and suspends associated accounts, bots, domains or subreddits, which the Responsible Builder Policy lists as its enforcement actions. Data API Terms 3.2 adds a permanent block on Data API access where Reddit believes you breached the limits clause. Because the breach is contractual, Reddit acts on its own authority and does not wait for a court.

What happens to my access if I close my Reddit account after scraping?

Closing the account does not retire the obligation. Reddit lists section 7, the clause banning scraping, among the sections that survive termination. The contract claim rests on the act of scraping, not on whether the account still exists.

How does Reddit decide who gets a data licensing deal?

Reddit's Public Content Policy frames licensing as a way to control access rather than a market open to anyone. Reddit enters these arrangements to know who is accessing public content and why, to place contractual restrictions on use, and to hold licensees to honoring deletions Redditors make. The policy names unauthorized scraping and data brokers as the behavior licensing is meant to replace.

Can university researchers get special Reddit data access?

Academic research runs through a separate channel from the standard Data API. Reddit routes it through the Reddit for Researchers program rather than general API approval or scraping. Scraping and unapproved API use stay off limits for research the same as for any other use case.

Can Reddit change Data API rate limits whenever it wants?

Reddit keeps that decision to itself. Data API Terms 2.9 gives Reddit the right to set and enforce limits on Data API use in its sole discretion, and section 3.2 bars developers from working around whatever limit is currently in force. A client built against today's limits can find itself throttled or blocked tomorrow with no negotiation.

See us more often in Google

One click marks RedReplier as a preferred source, so our articles sit higher in your Top Stories, AI Mode, and AI Overviews.

Before you go...

RedReplier

RedReplier

Catch every buyer asking for what you sell

RedReplier watches Reddit, X, Bluesky and Hacker News in real time, ranks every thread by buyer intent, and drafts your reply, so you get there first.

Reddit, X, Bluesky & HN

Real-time intent alerts

Unlimited AI replies

Ranked by buyer intent

Related Articles