This article is general information, not legal advice. Consult a qualified attorney for guidance specific to your situation.
Twenty state privacy laws are in effect. Zero comprehensive federal privacy statutes exist. Four separate platform-versus-scraper legal theories are in active litigation. That patchwork is the operative rulebook for US web scraping laws in 2026.
The map keeps moving. Indiana, Kentucky, and Rhode Island brought their laws live on January 1, 2026, and Connecticut’s amendments followed on July 1 (MultiState, 2026). For the global “is web scraping legal” question, see our web scraping legal compliance guide. This page answers the next question: which laws reach the data your team actually collects.
By the end, you will hold a current state-by-state map plus a use-case framework that says which statutes and cases apply to price data, people data, images, AI corpora, and everything behind a login.
Quick Digest
- The 2023 to 2026 shift: 5 state privacy laws effective at the end of 2023 became 20 by July 2026, state attorneys general began enforcing, and platform litigation moved past the CFAA.
- The federal layer: hiQ v. LinkedIn’s full arc is a CFAA win followed by a contract loss, a $500,000 consent judgment, and a permanent injunction (2022).
- The 20-state map: 20 comprehensive state privacy laws are in effect as of July 2026 (19 by IAPP’s stricter count), each with a “publicly available” carve-out narrower than most teams assume.
- Special-category data: faceprints from public photos cost Clearview AI 23% of its equity ($51.75M value, 2025); Texas v. Meta settled biometric claims at $1.4B (2024).
- Data-broker registration: selling scraped people-data can make you a broker in California, Texas, Vermont, and Oregon; California’s DROP deadline lands August 1, 2026, at $200 per-request per-day.
- The AI layer: Bartz v. Anthropic’s $1.5B settlement (final approval July 20, 2026) turned on data provenance, not training; fair use for AI training is unsettled on appeal.
- Enforcement in dollars: verified outcomes run from $345,178 (Todd Snyder, 2025) to $1.4B (Texas v. Meta, 2024); Healthline’s $1.55M is the largest CCPA settlement to date.
- The use-case map: exposure ranks by data category, access method, and downstream use, not page publicness; six rows map each pipeline to the laws it triggers.
- The compliance baseline: four gates before collection plus a vendor-diligence question set, because exposure does not transfer with the invoice.
Why US Scraping Compliance Changed Between 2023 and 2026
Three verified shifts define the 2026 landscape. First, the state-law wave: the first five comprehensive state privacy laws had all taken effect by December 31, 2023, and the count stands at 20 as of July 2026 (MultiState, 2026; IAPP counts 19 because Florida’s law is gated to $1B+ revenue companies). Second, enforcement became real: Texas filed the first-ever attorney general suit under a comprehensive state privacy law in January 2025, against Allstate and its data subsidiary Arity. Third, platform litigation moved past the CFAA (Computer Fraud and Abuse Act) to contract, copyright, and DMCA §1201 anti-circumvention claims.
The operative question for a data team changed shape with it: not “can we access this page” but “what does the extracted dataset contain, and where do its subjects live.” The state map matters this much because of an absence. The American Privacy Rights Act (APRA) expired with the 118th Congress in January 2025 and had not been reintroduced as of spring 2026 (R Street, 2026). Sectoral federal statutes (HIPAA, GLBA, COPPA, FCRA) still reach scraped data that touches their domains, but no general-purpose federal privacy statute sits above the states.
No comprehensive federal privacy law exists. The 20-state map below is the operative rulebook for scraped personal data.
Quick Summary
Q: Why did US web scraping compliance change so much between 2023 and 2026?
A: Three shifts: the count of effective state privacy laws quadrupled from 5 to 20 between late 2023 and July 2026, state attorneys general started enforcing (Texas v. Allstate, filed January 2025), and platforms swapped CFAA theories for contract, copyright, and anti-circumvention claims. With no federal statute, the state map became the rulebook.
The Federal Layer: CFAA, Copyright, DMCA §1201, and Terms of Service
Four federal theories actually reach scraping: the CFAA, copyright and fair use, DMCA §1201 anti-circumvention, and breach of contract through terms of service. Since 2024, platforms have largely stopped bringing CFAA claims against logged-off public scraping and shifted their weight to the other three (ZwillGen, 2026).

| Case | Court | What was held | Status (as of July 2026) |
|---|---|---|---|
| Van Buren v. United States | US Supreme Court, 2021 | CFAA narrowed to a “gates-up-or-down” access test | Final; controlling precedent |
| hiQ Labs v. LinkedIn | 9th Cir. / N.D. Cal., 2022 | CFAA win for scraper; contract loss on remand | Closed: $500K consent judgment + injunction |
| Meta v. Bright Data | N.D. Cal., Jan 2024 | Meta’s terms do not bar logged-off scraping of public data | District-level ruling; terms-specific |
| X Corp v. Bright Data | N.D. Cal., May 2024 | Claims dismissed; state claims preempted by Copyright Act | Settled confidentially June 2025; no merits precedent |
| Ryanair v. Booking | D. Del. / 3d Cir. | $5,000 CFAA jury verdict erased by JMOL, Jan 2025 | On appeal to the Third Circuit; pending |
| Reddit v. Anthropic | Cal. state court, filed June 2025 | Contract-only theory against AI-training scraping | Pending |
| Reddit v. Perplexity, SerpApi, Oxylabs | S.D.N.Y., filed Oct 2025 | DMCA §1201 theory for bypassing technical measures | Pending |
Sources: court records and law-firm analyses (Jenner & Block 2022; ZwillGen 2022, 2026; Farella Braun + Martel 2024; Cooley 2025). Table current as of July 2026.
The CFAA after Van Buren and hiQ (told completely)
In Van Buren v. United States (2021), the Supreme Court described CFAA liability as a “gates-up-or-down inquiry”: one “either can or cannot access a computer system,” and one “either can or cannot access certain areas within the system” (Supreme Court, 2021). Using access you legitimately hold for a purpose the owner dislikes is not a federal crime. Footnote 8 reserved whether the gates are set only by code or also by contracts.
hiQ Labs v. LinkedIn is the case teams half-remember, and the half they remember is the safe half. The Ninth Circuit held in April 2022 that scraping public LinkedIn pages likely does not violate the CFAA. The ending went the other way: in November 2022 the district court granted LinkedIn summary judgment on breach of contract, and the case closed with a $500,000 consent judgment, a permanent injunction, and deletion of all scraped data, code, and derived algorithms (ZwillGen, 2022).
hiQ’s CFAA win did not make terms of service irrelevant. On remand, hiQ lost LinkedIn’s breach-of-contract claim and ended under a $500,000 consent judgment and a permanent injunction.
Ryanair v. Booking collapsed on arithmetic. A Delaware jury found a CFAA violation over partly login-gated fare data in July 2024 and awarded exactly $5,000, the civil loss threshold; in January 2025 Judge William Bryson erased the verdict because only $2,457.72 of the claimed losses qualified as CFAA “loss” (Cooley, 2025). The appeal is pending at the Third Circuit as of July 2026. CFAA exposure concentrates behind logins and technical barriers, and even there the loss threshold is a litigated gate.
Copyright and fair use: content versus facts
Copyright protects scraped content, not scraped facts: Feist Publications v. Rural Telephone (Supreme Court, 1991) settled that facts and non-creative compilations carry no copyright. Fair use is a four-factor, case-by-case defense, not a safe harbor. The older trespass to chattels theory is still pleaded on occasion, but it requires proof of actual impairment to the server and has been mostly dormant in the modern cases.
The preemption point comes from X Corp v. Bright Data. Dismissing X’s claims in May 2024, Judge William Alsup wrote: “X Corp. wants it both ways: to keep its safe harbors yet exercise a copyright owner’s right to exclude, wresting fees from those who wish to extract and copy X users’ content” (CNBC, 2024). State-law claims over copying public data were held preempted by the Copyright Act. One caution: the parties settled confidentially in June 2025, so no final merits precedent exists; do not cite this case as “scrapers won.”
DMCA §1201: bypassing anti-bot measures is the new front
DMCA §1201 prohibits circumventing technical measures that control access to protected works. In October 2025, Reddit sued Perplexity AI, SerpApi, and Oxylabs in the Southern District of New York on this theory, over scraping Reddit content through Google’s search results (ZwillGen, 2026); the suit is pending as of July 2026. The practical read: bypassing rate limits, CAPTCHAs, or anti-bot systems now carries the highest theory-risk of any scraping behavior. robots.txt remains a norm rather than a law, but a documented pattern of ignoring it feeds these narratives in court.
Terms of Service: logged-in versus logged-off
Contract is the theory that actually ended hiQ, and it turns on whether you are bound: clickwrap terms bind account holders, browsewrap terms bind far less reliably. In Meta v. Bright Data (January 2024), Judge Edward Chen granted the scraping company summary judgment: “The Facebook and Instagram Terms do not bar logged-off scraping of public data; perforce it does not prohibit the sale of such public data” (Farella Braun + Martel, 2024). Non-account-holders, the court noted, do not see the terms and cannot be bound. Two qualifications: it is a district-court ruling tied to Meta’s specific terms, and Reddit v. Anthropic (filed June 2025, pending as of July 2026) is testing a broader contract-only theory. Terms exposure is civil, not criminal; as hiQ shows, civil is enough to end a data business. That is the federal half of the map. The other half is written state by state.
Quick Summary
Q: Which federal laws actually apply to web scraping in the US?
A: Four theories: the CFAA (narrow after Van Buren and hiQ for public, logged-off data), copyright (content, not facts, with preemption of state-law workarounds), DMCA §1201 (bypassing anti-bot measures, the emerging front), and terms-of-service contract claims (the theory that ended hiQ with a $500,000 judgment). Access method decides which theory bites.
Which States Regulate Scraped Personal Data in 2026?
Twenty states have comprehensive consumer privacy laws in effect as of July 2026; IAPP’s stricter count says 19, excluding Florida because its law applies only to companies with $1B+ global revenue. The 2026 wave alone brought three new laws live on January 1 (Indiana, Kentucky, Rhode Island), Connecticut’s SB 1295 amendments on July 1, and new California data-broker requirements on August 1. The map moves 2 to 3 times a year, hence the review stamp on the table.
A scraping team reads the table with three questions per state: does our dataset contain personal data of that state’s residents, do we trip the applicability threshold, and does the “publicly available” carve-out actually cover what we collected?
The 2026 map: 20 laws, one table
| State | Statute | Effective date | Carve-out model | Scraping-relevant notes |
|---|---|---|---|---|
| California | CCPA (amended by CPRA) | Jan 1, 2020; CPRA amendments Jan 1, 2023 | CA-model (narrower) | Revenue OR volume thresholds; biometric exclusion; broker registry; ADMT regs eff. Jan 1, 2026 |
| Virginia | VCDPA | Jan 1, 2023 | VA-model | The template most later states copied; opt-in consent for sensitive data |
| Colorado | CPA | Jul 1, 2023 | VA-model | Recognizes universal opt-out signals (GPC) |
| Connecticut | CTDPA + SB 1295 | Jul 1, 2023; amendments Jul 1, 2026 | VA-model | 35,000-consumer threshold; ANY sensitive-data processing or sale triggers scope |
| Utah | UCPA | Dec 31, 2023 | VA-model | Business-friendly variant; note the odd effective date |
| Texas | TDPSA | Jul 1, 2024 | VA-model | No volume threshold; data-broker registry; first AG enforcement (Allstate) |
| Oregon | OCPA | Jul 1, 2024 | VA-model | Broker registry; further provisions phased through Jan 1, 2026 |
| Florida | FDBR | Jul 1, 2024 | VA-model | $1B+ revenue gate; the reason IAPP counts 19 |
| Montana | MCDPA | Oct 1, 2024 | VA-model | Standard VA-model mechanics |
| Delaware | DPDPA | Jan 1, 2025 | VA-model | Standard VA-model mechanics |
| Iowa | ICDPA | Jan 1, 2025 | VA-model | Narrower rights set than most peers |
| Nebraska | NDPA | Jan 1, 2025 | VA-model | TDPSA-style structure |
| New Hampshire | NHDPA | Jan 1, 2025 | VA-model | Standard VA-model mechanics |
| New Jersey | NJDPA | Jan 15, 2025 | VA-model | Note the mid-month effective date |
| Tennessee | TIPA | Jul 1, 2025 | VA-model | Standard VA-model mechanics |
| Minnesota | MCDPA | Jul 31, 2025 | VA-model | Standard VA-model mechanics |
| Maryland | MODPA | Oct 1, 2025 | VA-model text, strictest duties | Strict-necessity minimization; flat ban on selling sensitive data |
| Indiana | INCDPA | Jan 1, 2026 | VA-model | 2026 newcomer |
| Kentucky | KCDPA | Jan 1, 2026 | VA-model | 2026 newcomer |
| Rhode Island | RIDTPPA | Jan 1, 2026 | VA-model | Low 35,000-consumer thresholds (MultiState, 2026) |
Sources: MultiState (Feb 2026), IAPP (Jan 2026), and the IAPP US State Privacy Legislation Tracker. Last reviewed July 2026.

What “publicly available” actually exempts (narrower than scrapers assume)
Every one of the 20 laws excludes “publicly available information” from its definition of personal data, and the definitions differ. Keep two terms distinct: publicly accessible describes a page anyone can load; “publicly available” is a statutory term of art with conditions attached.
California’s is the narrower model. Under Cal. Civ. Code § 1798.140(v)(2), “publicly available” means: “(I) Information that is lawfully made available from federal, state, or local government records. (II) Information that a business has a reasonable basis to believe is lawfully made available to the general public by the consumer or from widely distributed media. (III) Information made available by a person to whom the consumer has disclosed the information if the consumer has not restricted the information to a specific audience.” The statute then subtracts: “‘Publicly available’ does not mean biometric information collected by a business about a consumer without the consumer’s knowledge.”
The Virginia-model states run broader but keep the same limit: content the consumer restricted to a specific audience does not qualify. Logged-in, followers-only, or privacy-limited content fails the carve-out in every state.
Public on the internet does not equal “publicly available” under privacy law. California expressly excludes biometric data collected without the consumer’s knowledge, and every state excludes content the consumer restricted to a specific audience.

One obligation attaches regardless of the carve-out analysis: deletion and opt-out duties bind the data holder however the data was gathered. Re-scraping a record a consumer already deleted has no clean statutory answer as of July 2026; treat it as an open operational question.
CCPA/CPRA deep dive: the California stack
The CCPA applies at $25M+ annual revenue, or personal information of 100,000+ consumers or households bought, sold, or shared, or 50%+ of revenue from selling or sharing personal information (IAPP, 2026). Scraped personal information triggers notice-at-collection duties, and “sale” is broader than cash: any transfer for “monetary or other valuable consideration” counts, a reading the California AG established against Sephora in 2022. California also recognizes the Global Privacy Control (GPC) opt-out signal, and its ADMT regulations took effect January 1, 2026.
Penalties run $2,500 per violation and $7,500 per intentional violation or violations involving minors’ data, and the CPPA’s Honda fine multiplied affected consumers by $2,500: per-consumer stacking is regulator practice, not a theoretical ceiling (Cal. Civ. Code § 1798.155; CPPA, 2025). The Delete Act’s broker mechanics get their own section below.
Texas TDPSA and the first-ever state enforcement (Texas v. Allstate/Arity)
The TDPSA has no volume threshold at all: any non-small-business processing Texans’ personal data is in scope, and even SBA-defined small businesses need consent to sell sensitive data (IAPP, 2026).
Then Texas used it. On January 13, 2025, the Texas Attorney General sued Allstate and its data subsidiary Arity, the first AG suit under any comprehensive state privacy law, alleging covert collection and sale of driving data covering 45M+ consumers and seeking up to $7,500 per violation under the TDPSA plus $10,000 per violation under the broker law. The case is pending as of July 2026: the filing and the penalties sought are the facts, no outcome exists. The signal is still legible: covert large-scale collection plus resale is now an attorney general priority.
The strict outliers: Maryland, Connecticut, and the minors’ rules
Maryland’s MODPA (effective October 1, 2025, applying to processing from April 1, 2026) rewrote the default: collection is allowed only as “reasonably necessary and proportionate” to a product or service “requested by the consumer,” a duty rather than a disclosure. A scraped dataset of Maryland residents was, by construction, not collected to provide something the consumer requested; commentators read MODPA as the state law most structurally hostile to bulk collection. It also carries the first flat ban on selling sensitive data; no consent cures it (EPIC, 2025).
Connecticut’s SB 1295 (effective July 1, 2026) cut the threshold from 100,000 to 35,000 consumers and extended the law to any entity processing sensitive data or selling personal data, regardless of volume (Covington, 2025). Even a small number of Connecticut residents’ sensitive records can now pull an out-of-state company into scope. Oregon’s minors’ amendments and Rhode Island’s low 35,000-consumer thresholds round out the tightening trend. And where the carve-outs subtract a category outright, the special-category statutes below take over.
Quick Summary
Q: Which US states regulate scraped personal data in 2026?
A: Twenty states have comprehensive privacy laws in effect as of July 2026 (19 by IAPP’s stricter count). All exempt “publicly available” data but define it narrowly: audience-restricted content never qualifies, and California excludes covertly collected biometrics outright. The strict outliers, Maryland’s necessity standard, Connecticut’s no-threshold sensitive-data trigger, and Texas’s no-volume-threshold law, are where scraped datasets most often land in scope.
Special-Category Data: Biometrics, Health, and Minors
Scraping liability concentrates by data category, not by page accessibility. Three categories have already produced nine-to-ten-figure outcomes or carry private rights of action.
Biometrics. Illinois’s BIPA (740 ILCS 14) requires informed written consent before capturing biometric identifiers, including faceprints computed from photos, with a private right of action at $1,000 per negligent and $5,000 per intentional or reckless violation (a 2024 amendment, SB 2979, made repeated scans of one person a single violation). Clearview AI is the priced example: a faceprint database built from reportedly 60+ billion public images ended in an Illinois class settlement, final-approved by Judge Sharon Johnson Coleman on March 20, 2025, valued at $51.75 million and paid as a 23% equity stake because the company lacked the cash; 22 state attorneys general and DC opposed it (Troutman, 2025). A 2022 ACLU consent decree separately bans Clearview from selling its faceprint database to most private businesses. Texas’s CUBI, AG-enforced at up to $25,000 per violation, produced Texas v. Meta: $1.4 billion on July 30, 2024, the largest single-state AG privacy settlement at the time (Texas AG, 2024). Washington’s HB 1493 (2017) is the third biometric statute. The screen: does the pipeline touch faces or photos convertible to faceprints?
Health-inferable data. Washington’s My Health My Data Act (in force March 31, 2024) regulates consumer health data outside HIPAA, broadly enough to reach inferences from searches, purchases, and location near clinics. It carries a private right of action via Washington’s Consumer Protection Act: actual damages plus treble damages up to $25,000. The first class action, Maxwell v. Amazon, was filed February 10, 2025 and is pending as of July 2026 (Orrick, 2025). Nevada’s SB 370 covers similar ground without the private right of action. The screen: “we don’t collect health data” fails if health status is inferable from what you do collect.
Minors’ data. Known-minor audiences carry sale and targeted-advertising bans across the newer state laws, with age-appropriate design codes spreading through 2026 amendments (Connecticut, Oregon). Adjacent regime, one line: names and likenesses in commercial products can trigger state rights of publicity independent of any privacy statute.
The through-line: public photos are not consent. Data category is one trigger; what a team does with the data downstream is the next.
Quick Summary
Q: Why are biometrics, health data, and minors’ data the highest-risk categories for scraping?
A: They carry consent requirements and private rights of action that public accessibility does not cure. Clearview AI’s faceprints from public photos settled at 23% of company equity ($51.75M value, 2025); Texas v. Meta settled biometric claims at $1.4B (2024); Washington’s My Health My Data Act lets consumers sue over inferred health data, with treble damages up to $25,000.
Data-Broker Registration: When Selling Scraped People-Data Makes You a Broker
Selling scraped personal data about consumers you have no direct relationship with presumptively makes you a data broker in the four registry states: California, Texas, Vermont, and Oregon. Registration is annual and separate from the privacy laws above.
California added teeth. The Delete Act (SB 362) built DROP, the CPPA’s one-stop deletion platform: registered brokers had to begin accessing it January 1, 2026 and must process all deletion requests by August 1, 2026, days after this article’s review date, at $200 per-deletion-request per-day penalties that compound across consumers (IAPP, 2026). It is the single most concrete compliance date in this article’s scope.
Registration failure is also the easy count for an enforcer: Texas charged Allstate with failure to register alongside its TDPSA counts, seeking $10,000 per violation on registration alone (Texas AG, 2025, pending). One caveat: “we only sell aggregated or derived data” does not automatically exit broker definitions; the statutory wording governs per state.
Quick Summary
Q: Do you have to register as a data broker if you sell scraped people-data?
A: In California, Texas, Vermont, and Oregon, yes, presumptively, if the consumers are ones you have no direct relationship with. California’s DROP deletion deadline of August 1, 2026 carries $200 per-request per-day penalties, and Texas already paired a $10,000-per-violation registration count with its first privacy enforcement suit.
The AI Layer: Training-Data Rulings and the 2026 AI Statutes
Reselling scraped data is one regulated downstream use. Training AI on it is the other. Two developments converged across 2025 and 2026: courts began ruling on fair use for AI training, and the first state AI statutes touching training data took effect January 1, 2026. The through-line of both is provenance.
The rulings. Bartz v. Anthropic split the question in two. In June 2025, Judge Alsup held that training LLMs on lawfully acquired books is fair use, but downloading from pirate libraries such as LibGen was not protected: the acquisition channel, not the training, created the liability. The settlement priced that channel: final approval on July 20, 2026 at $1.5 billion, the largest copyright class settlement in US history, roughly $3,000 per work across about 500,000 pirated books, with destruction of the pirated files ordered (Authors Alliance, 2026). Thomson Reuters v. Ross Intelligence points the other way: fair use rejected outright in February 2025, with the Third Circuit hearing argument June 11, 2026, the first federal appellate case squarely on the question. No decision exists as of July 2026. NYT v. OpenAI remains active with no merits ruling as of July 2026; its January 2026 order compelling production of 20 million de-identified ChatGPT logs shows courts treating AI pipelines as fully discoverable (Pillsbury, 2026).
“Fair use covers AI training” is not a safe assumption: one court blessed lawful copies, another rejected fair use outright, and the appellate answer is pending. The answerable question is narrower: can you document where every record came from? Teams sourcing corpora can start with our guide to compliant web scraping for AI training data.
The statutes. California’s AB 2013 requires generative AI developers to publish training-data documentation for models made available in California, effective January 1, 2026; scraped corpora are squarely in scope (King & Spalding, 2025). Texas’s TRAIGA took effect the same day, intent-based with a NIST AI RMF safe harbor. Colorado repealed its AI Act for the SB 26-189 ADMT law, signed May 14, 2026, effective January 1, 2027 (Miller Nash, 2026). Documentation duties are arriving faster than the fair-use question is resolving.

Quick Summary
Q: Is it legal to scrape data for AI training in the US?
A: Unsettled, and provenance-driven. Bartz v. Anthropic blessed training on lawfully acquired copies but priced pirated acquisition at $1.5B (final approval July 20, 2026); Thomson Reuters v. Ross rejected fair use, with the first appellate ruling pending as of July 2026; and California’s AB 2013 requires GenAI developers to publish training-data documentation from January 1, 2026. The defensible position is an acquisition trail for every record.
Enforcement: What Non-Compliance Has Actually Cost
Enforcement moved from theory to a priced reality across 2024 and 2025. The enforcers are state attorneys general, California’s CPPA, the FTC, and private classes where statutes allow them (BIPA, MHMD).
| Action | Enforcer | Amount | Date | Why it matters for scraping teams |
|---|---|---|---|---|
| Sephora (CCPA) | CA AG | $1.2M | Aug 2022 | Historical first CCPA settlement; adtech data flows count as “sale” |
| DoorDash (CCPA/CalOPPA) | CA AG | $375K | Feb 2024 | Marketing co-op data swap treated as a “sale” without notice |
| American Honda (CCPA) | CPPA | $632,500 | Mar 2025 | First major CPPA fine; consumers × $2,500 stacking math |
| Todd Snyder (CCPA) | CPPA | $345,178 | May 2025 | Broken opt-out mechanism; over-collection during rights requests |
| Healthline (CCPA) | CA AG | $1.55M | Jul 2025 | Largest CCPA settlement to date; article-reading data implied diagnoses |
| Texas v. Meta (CUBI) | TX AG | $1.4B | Jul 2024 | Biometrics without consent; largest single-state AG privacy settlement at the time |
| Texas v. Allstate/Arity (TDPSA + broker law) | TX AG | Pending; seeks $7,500/violation + $10K/violation | Filed Jan 2025 | First AG suit under a comprehensive state privacy law; covert bulk collection |
| Clearview AI (BIPA class) | Private class | $51.75M value (23% equity) | Mar 2025 | Faceprints from public photos; equity-funded settlement |
| FTC: X-Mode & InMarket | FTC | Conduct bans | Jan 2024 | Banned from selling precise/sensitive location data |
| FTC: Gravy Analytics & Mobilewalla | FTC | Conduct bans | Dec 2024 | Third-party data without verified consent treated as unfair practice |
Sources: agency releases and law-firm alerts (CA AG 2022–2025; CPPA 2025; Texas AG 2024–2025; FTC 2024; Coblentz 2025; Cooley 2025). Table current as of July 2026.

Healthline’s $1.55M is the anchor for CCPA exposure: the largest settlement under the statute to date, over sharing article-reading data that implied medical diagnoses (CA AG, July 2025). The compounding mechanism sits in the Honda action: the CPPA computed the fine as affected consumers times $2,500. Small per-violation figures become large totals through ordinary regulator arithmetic.
The pattern across the dataset is consistent: sensitive and health-adjacent data (Healthline), biometrics (Texas v. Meta, Clearview), and covert bulk collection plus resale (Allstate, pending) draw the enforcement, not garden-variety public-page scraping. Two reading notes: Sephora is 2022 context, not current practice, and Allstate is a filing with penalties sought, not an outcome.
Quick Summary
Q: What has privacy non-compliance actually cost data collectors?
A: Verified 2024–2025 outcomes range from $345,178 (Todd Snyder, CPPA) to $1.4B (Texas v. Meta, biometrics), with Healthline’s $1.55M the largest CCPA settlement to date. Sensitive categories, biometrics, and covert bulk collection draw the enforcement, and per-violation stacking (Honda: consumers × $2,500) turns small statutory figures into large totals.
What Does Each Scraping Use Case Actually Trigger?
Same pipeline, different data, different legal surface. The map runs on three screening questions per project: does the extracted dataset contain personal data, and of which states’ residents; how is the source accessed (public logged-off, credentialed, or technically gated); and what is the downstream use (internal analytics, resale, AI training, commercial product)?
Risk ranks by data category, access method, and use. It does not rank by how public the page looks. That is the table’s whole argument.
| Use case | Laws/theories implicated | Realistic risk posture | Required controls |
|---|---|---|---|
| Price, product, and market data | Copyright thin (facts, per Feist); CFAA low if logged-off; ToS if logged-in | Lowest-risk row | Logged-off access; no content republishing |
| Company/firmographic data | Mostly non-personal; sole-proprietor and contact details edge into state laws | Low, with a personal-data edge | Field minimization; threshold awareness |
| People data (contacts, profiles, lead lists) | All 20 state laws; carve-out limits; “sale” breadth; broker registration if resold | Highest statutory surface | Rights-request and GPC handling; broker analysis; carve-out check per state |
| Images, media, and faces | Copyright (content); BIPA/CUBI/HB 1493 if faceprints; publicity rights if commercial | Highest severity per record | No faceprint computation without consent posture; licensing review |
| AI training corpora | Copyright/fair-use split (Bartz vs. Ross); AB 2013 documentation; provenance duty | Unsettled; provenance-driven | Acquisition-channel logging per record |
| Anything behind a login or paywall | ToS contract exposure (hiQ endgame); carve-outs fail; §1201 if barriers bypassed | Highest behavioral risk | Counsel sign-off before credentialed collection; no barrier circumvention |
Sources: consolidated from the cases and statutes above (court records, state statutes, agency releases). Stamped as of July 2026.

Row 1, price and market data, is the lowest-risk configuration: facts carry no copyright under Feist, and logged-off access keeps the CFAA and contract theories quiet. The posture holds only while it stays logged-off and republishes nothing.
Row 2, firmographic data, sits mostly outside the privacy statutes because companies are not consumers. The edge case: sole proprietors’ and named contacts’ details are personal data of the states they live in. Our firmographic data guide covers this category.
Row 3, people data, carries the most statutory surface: all 20 state laws, the narrow carve-outs, “sale” as any valuable consideration, and broker registration if resold. Threshold monitoring and rights handling stop being optional here.
Row 4, images and faces, is where severity per record peaks. A photo is content under copyright; a faceprint computed from it is a biometric identifier under BIPA, CUBI, and Washington’s HB 1493. Clearview’s 23%-equity settlement is this row read wrong (National Law Review, 2025).
Row 5, AI corpora, turns on provenance: Bartz priced pirated acquisition at $1.5B while blessing lawful copies, and AB 2013 requires the documentation regardless (Authors Alliance, 2026). The pending Third Circuit ruling in Ross could shift this row.
Row 6, logged-in or paywalled content, is where “public data” intuitions fail hardest: you accepted the terms, so contract exposure is live (the hiQ endgame); the carve-outs exclude audience-restricted content in every state; and bypassing barriers adds the §1201 theory. Judge Chen’s logged-off holding is the mirror image: the protection he described stops at the login wall (Farella Braun + Martel, 2024).
Quick Summary
Q: Which laws apply to each web scraping use case?
A: Price and product data implicates the least (uncopyrightable facts, logged-off access). People data implicates the most statutory surface (20 state laws, “sale” breadth, broker registration). Faces, AI corpora, and logged-in content each carry a dominant theory: biometric statutes, copyright plus provenance duties, and contract respectively. Risk follows data category, access method, and downstream use, not page publicness.
A Compliance Baseline for Scraping Teams and Scraped-Data Buyers
Everything above compresses into a program baseline with two halves: controls for teams that scrape, and diligence for teams that buy. Exposure does not transfer with the invoice.
Program controls for teams that scrape
The four gates, run before any collection starts:
- Robots/ToS posture: documented per source, on file, not in someone’s head.
- Login state: logged-off public-web collection only, unless counsel signs off on a credentialed exception.
- Technical barriers: no circumvention of anti-bot measures, rate limits, or CAPTCHAs.
- Data category: PII, biometric, health-inferable, and minors’ data flagged before collection, not after delivery.
Behind the gates sit the standing obligations: field-level minimization at extraction time, threshold monitoring by state, rights-request and GPC handling, provenance logging per record, and an annual law-tracking pass against the IAPP tracker (linked in the state table above), because the statutes move 2 to 3 times a year. Validation of what actually arrives belongs in the same loop; our data quality framework covers that layer.

Vendor diligence for teams that buy scraped data
“Our vendor handles compliance” is not a control. Obligations attach to the data holder, and the FTC’s Mobilewalla and Gravy Analytics orders sanctioned exactly this posture: relying on third-party-sourced data without verified consent (FTC, 2024). What a buyer controls is the diligence conversation. Six questions for any scraping vendor:
- Do you collect logged-off public data only?
- How do you document provenance per record?
- How do you handle deletion and opt-out requests downstream?
- Which sources’ ToS positions have you assessed, and where are they documented?
- What do you filter out at extraction time (fields, categories, minors’ signals)?
- What contractual warranties do you offer on acquisition method?
These are the questions Forage AI invites from its own prospects as a managed extraction provider; a vendor that cannot answer them in writing has answered them anyway. As a managed service, Forage operates compliance-aware extraction practices along the lines this article maps: a logged-off public-web collection boundary, field-level minimization at extraction time, and provenance tracking per record. The contract follows from the questions: acquisition-method warranties, provenance documentation, deletion-passthrough terms. For what a managed provider should own end to end, see what managed web data extraction covers.

Quick Summary
Q: What does a defensible compliance baseline for web scraping look like?
A: Four gates before collection (documented robots/ToS posture, logged-off-only access, no barrier circumvention, a data-category screen) and standing obligations after it (field minimization, threshold monitoring by state, rights-request and GPC handling, provenance logging, annual law tracking). Buyers add a vendor-diligence question set and contract terms on acquisition method, because exposure does not transfer with the invoice.
Expert Insights
The strongest lines in the map above belong to the judges, statutes, and regulators; they are collected here.
Expert Insights
“X Corp. wants it both ways: to keep its safe harbors yet exercise a copyright owner’s right to exclude, wresting fees from those who wish to extract and copy X users’ content,” wrote Judge Alsup, adding that “free rein” over public web data “risks the possible creation of information monopolies that would disserve the public interest.” (Judge William Alsup, N.D. Cal., X Corp v. Bright Data, May 2024)
“If this is what passes as technological harm, then more CFAA absurdity is certain to come in the near future.” (Kieran McCarthy, founding partner, McCarthy Law Group, on Ryanair v. Booking; Technology & Marketing Law Blog, March 2025)
The statutory texts carry the argument. California’s carve-out has a built-in subtraction: “‘Publicly available’ does not mean biometric information collected by a business about a consumer without the consumer’s knowledge” (Cal. Civ. Code § 1798.140(v)(2)). Maryland inverts the model, permitting collection only when “reasonably necessary and proportionate” to a service “requested by the consumer” (MODPA, per EPIC, 2025). The two mark the edges of the 20-state range.
Three anchors price the heavy rows of the use-case map: Judge Chen’s holding that Meta’s terms “do not bar logged-off scraping of public data” bounds row 6 from the safe side (N.D. Cal., January 2024); Clearview’s $51.75M equity settlement prices row 4 (final approval March 2025); Bartz’s $1.5B settlement over pirated acquisition prices row 5 (final approval July 2026).
The FTC’s December 2024 orders against Mobilewalla and Gravy Analytics are the closest regulatory statement on scraped-data supply chains: acquiring location data from third parties without verifying consumer consent was treated as an unfair practice, and Mobilewalla’s order carried the first-ever prohibition on collecting data from real-time bidding streams (FTC, 2024). Buyers, not only collectors, sit inside the enforcement perimeter.
Frequently Asked Questions
Which US states have privacy laws that affect web scraping?
Twenty states have comprehensive privacy laws in effect as of July 2026; IAPP’s stricter count says 19, excluding Florida’s $1B+ revenue-gated law. The newest wave brought Indiana, Kentucky, and Rhode Island live on January 1, 2026, with Connecticut’s amendments following July 1. The full effective-date table sits in the state-layer section above.
Does the CFAA make scraping public data illegal?
No, not for public, logged-off data in the Ninth Circuit, under Van Buren’s gates-up-or-down test and hiQ. The catch: hiQ still lost on breach of contract and closed under a $500,000 consent judgment. CFAA claims persist for credentialed and gated access; the Ryanair appeal is pending at the Third Circuit as of July 2026.
Is publicly available data exempt from state privacy laws?
Only within each statute’s definition, which is narrower than the everyday meaning. Content the consumer restricted to a specific audience never qualifies, and California expressly excludes biometric data collected without the consumer’s knowledge. A page being loadable by anyone does not make its data “publicly available” in the statutory sense.
Can you scrape faces or photos in the US?
Photos as content raise copyright questions; converting them into faceprints triggers the biometric statutes (Illinois BIPA, Texas CUBI, Washington HB 1493), with consent requirements and, in Illinois, a private right of action. Clearview AI scraped only publicly posted images and its Illinois settlement was valued at 23% of the company’s equity. Public posting is not consent.
What penalties apply if a scraping program violates the CCPA?
The statute sets $2,500 per violation and $7,500 per intentional violation or violations involving minors’ data, and regulators stack these per consumer, as the Honda fine’s consumers-times-$2,500 arithmetic showed in 2025. Healthline’s $1.55M settlement (July 2025) is the current ceiling.
Do Terms of Service override the right to scrape public pages?
Terms bind those who accept them. Meta’s terms were held not to bar logged-off scraping of public data (Meta v. Bright Data, 2024), but logged-in scraping means you accepted the terms, and that contract exposure is how hiQ ended. Reddit v. Anthropic, pending as of July 2026, is testing how far a contract-only theory stretches.
Conclusion: The Map Is the Law
With no comprehensive federal statute, the operative US rulebook for scraped data is the map you just worked through: 20 state laws with their carve-outs and thresholds, four federal theories sorted by access method, and a use-case grid that ties them to the data your team actually collects. That was the promise at the top of this page, and it is now your working artifact: scope each project against the three screening questions, the four gates, and the state table.
Four pending matters can shift this page: Thomson Reuters v. Ross and Ryanair v. Booking at the Third Circuit, summary judgment in NYT v. OpenAI, and California’s DROP deadline on August 1, 2026. This page is reviewed against the IAPP tracker; last reviewed July 2026.
Teams that would rather have a managed provider own this operational surface can start with the vendor-diligence questions above, and with the build-versus-buy decision guide for the wider sourcing decision.
Related Articles
- Web Scraping Legal Compliance Guide – The global “is web scraping legal” question: CFAA basics, GDPR and CCPA principles, and ethics.
- Web Data Extraction: Build vs. Buy Decision Guide – The sourcing decision that sits underneath the vendor-diligence questions.
- What Is Managed Web Data Extraction? – What a managed provider should own end to end, including the compliance-adjacent layers.
- A Data Quality Framework for External Data – The QA and validation layer that pairs with the program controls above.
Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.