Compliance & Regulation in Data Extraction

US Web Scraping Laws in 2026: State Privacy Laws, Federal Law, and a Use-Case Map for Data Teams

July 24, 2026

5 min read


Sai S

US Web Scraping Laws in 2026: State Privacy Laws, Federal Law, and a Use-Case Map for Data Teams featured image

This article is general information, not legal advice. Consult a qualified attorney for guidance specific to your situation.

Twenty state privacy laws are in effect. Zero comprehensive federal privacy statutes exist. Four separate platform-versus-scraper legal theories are in active litigation. That patchwork is the operative rulebook for US web scraping laws in 2026.

The map keeps moving. Indiana, Kentucky, and Rhode Island brought their laws live on January 1, 2026, and Connecticut’s amendments followed on July 1 (MultiState, 2026). For the global “is web scraping legal” question, see our web scraping legal compliance guide. This page answers the next question: which laws reach the data your team actually collects.

By the end, you will hold a current state-by-state map plus a use-case framework that says which statutes and cases apply to price data, people data, images, AI corpora, and everything behind a login.

Quick Digest

  • The 2023 to 2026 shift: 5 state privacy laws effective at the end of 2023 became 20 by July 2026, state attorneys general began enforcing, and platform litigation moved past the CFAA.
  • The federal layer: hiQ v. LinkedIn’s full arc is a CFAA win followed by a contract loss, a $500,000 consent judgment, and a permanent injunction (2022).
  • The 20-state map: 20 comprehensive state privacy laws are in effect as of July 2026 (19 by IAPP’s stricter count), each with a “publicly available” carve-out narrower than most teams assume.
  • Special-category data: faceprints from public photos cost Clearview AI 23% of its equity ($51.75M value, 2025); Texas v. Meta settled biometric claims at $1.4B (2024).
  • Data-broker registration: selling scraped people-data can make you a broker in California, Texas, Vermont, and Oregon; California’s DROP deadline lands August 1, 2026, at $200 per-request per-day.
  • The AI layer: Bartz v. Anthropic’s $1.5B settlement (final approval July 20, 2026) turned on data provenance, not training; fair use for AI training is unsettled on appeal.
  • Enforcement in dollars: verified outcomes run from $345,178 (Todd Snyder, 2025) to $1.4B (Texas v. Meta, 2024); Healthline’s $1.55M is the largest CCPA settlement to date.
  • The use-case map: exposure ranks by data category, access method, and downstream use, not page publicness; six rows map each pipeline to the laws it triggers.
  • The compliance baseline: four gates before collection plus a vendor-diligence question set, because exposure does not transfer with the invoice.
2026 Edition · Strategic Guide
How to Get Started With Your Data Acquisition Strategy For AI
A strategic guide for data leaders who don’t know where to start.
Most guides about data infrastructure jump to the technical fix. This one starts a step earlier, at the strategy decision. It helps you see where you stand on the data acquisition maturity curve, what your options are, and what to ask before you pick a partner.
5 Data Acquisition Stages
3 Data Solutions
15 Min Read
Download the e-book
Free. Sent straight to your inbox.
We’ll email you the guide. No spam, unsubscribe anytime.

Why US Scraping Compliance Changed Between 2023 and 2026

Three verified shifts define the 2026 landscape. First, the state-law wave: the first five comprehensive state privacy laws had all taken effect by December 31, 2023, and the count stands at 20 as of July 2026 (MultiState, 2026; IAPP counts 19 because Florida’s law is gated to $1B+ revenue companies). Second, enforcement became real: Texas filed the first-ever attorney general suit under a comprehensive state privacy law in January 2025, against Allstate and its data subsidiary Arity. Third, platform litigation moved past the CFAA (Computer Fraud and Abuse Act) to contract, copyright, and DMCA §1201 anti-circumvention claims.

The operative question for a data team changed shape with it: not “can we access this page” but “what does the extracted dataset contain, and where do its subjects live.” The state map matters this much because of an absence. The American Privacy Rights Act (APRA) expired with the 118th Congress in January 2025 and had not been reintroduced as of spring 2026 (R Street, 2026). Sectoral federal statutes (HIPAA, GLBA, COPPA, FCRA) still reach scraped data that touches their domains, but no general-purpose federal privacy statute sits above the states.

No comprehensive federal privacy law exists. The 20-state map below is the operative rulebook for scraped personal data.

Quick Summary

Q: Why did US web scraping compliance change so much between 2023 and 2026?

A: Three shifts: the count of effective state privacy laws quadrupled from 5 to 20 between late 2023 and July 2026, state attorneys general started enforcing (Texas v. Allstate, filed January 2025), and platforms swapped CFAA theories for contract, copyright, and anti-circumvention claims. With no federal statute, the state map became the rulebook.

The Federal Layer: CFAA, Copyright, DMCA §1201, and Terms of Service

Four federal theories actually reach scraping: the CFAA, copyright and fair use, DMCA §1201 anti-circumvention, and breach of contract through terms of service. Since 2024, platforms have largely stopped bringing CFAA claims against logged-off public scraping and shifted their weight to the other three (ZwillGen, 2026).

The four federal legal theories that reach web scraping, with the litigation shift away from the CFAA. One, the Computer Fraud and Abuse Act, narrowed by Van Buren's 2021 gates-up-or-down test, with hiQ's April 2022 win covering public logged-off pages. Two, terms of service contract claims, the theory that actually ended hiQ: a November 2022 contract loss, a $500,000 consent judgment, and a permanent injunction. Three, copyright and fair use: Feist (1991) holds facts carry no copyright, and X Corp v. Bright Data's state-law claims were held preempted in 2024 before a confidential 2025 settlement. Four, DMCA section 1201 anti-circumvention, the new front: Reddit sued Perplexity, SerpApi, and Oxylabs in October 2025 over bot-barrier bypass, pending as of July 2026. Sources: court records and law-firm analyses cited in the article (Jenner & Block 2022; ZwillGen 2022, 2026; Cooley 2025).
Four theories, and the weight has moved off the CFAA
CaseCourtWhat was heldStatus (as of July 2026)
Van Buren v. United StatesUS Supreme Court, 2021CFAA narrowed to a “gates-up-or-down” access testFinal; controlling precedent
hiQ Labs v. LinkedIn9th Cir. / N.D. Cal., 2022CFAA win for scraper; contract loss on remandClosed: $500K consent judgment + injunction
Meta v. Bright DataN.D. Cal., Jan 2024Meta’s terms do not bar logged-off scraping of public dataDistrict-level ruling; terms-specific
X Corp v. Bright DataN.D. Cal., May 2024Claims dismissed; state claims preempted by Copyright ActSettled confidentially June 2025; no merits precedent
Ryanair v. BookingD. Del. / 3d Cir.$5,000 CFAA jury verdict erased by JMOL, Jan 2025On appeal to the Third Circuit; pending
Reddit v. AnthropicCal. state court, filed June 2025Contract-only theory against AI-training scrapingPending
Reddit v. Perplexity, SerpApi, OxylabsS.D.N.Y., filed Oct 2025DMCA §1201 theory for bypassing technical measuresPending

Sources: court records and law-firm analyses (Jenner & Block 2022; ZwillGen 2022, 2026; Farella Braun + Martel 2024; Cooley 2025). Table current as of July 2026.

The CFAA after Van Buren and hiQ (told completely)

In Van Buren v. United States (2021), the Supreme Court described CFAA liability as a “gates-up-or-down inquiry”: one “either can or cannot access a computer system,” and one “either can or cannot access certain areas within the system” (Supreme Court, 2021). Using access you legitimately hold for a purpose the owner dislikes is not a federal crime. Footnote 8 reserved whether the gates are set only by code or also by contracts.

hiQ Labs v. LinkedIn is the case teams half-remember, and the half they remember is the safe half. The Ninth Circuit held in April 2022 that scraping public LinkedIn pages likely does not violate the CFAA. The ending went the other way: in November 2022 the district court granted LinkedIn summary judgment on breach of contract, and the case closed with a $500,000 consent judgment, a permanent injunction, and deletion of all scraped data, code, and derived algorithms (ZwillGen, 2022).

hiQ’s CFAA win did not make terms of service irrelevant. On remand, hiQ lost LinkedIn’s breach-of-contract claim and ended under a $500,000 consent judgment and a permanent injunction.

Ryanair v. Booking collapsed on arithmetic. A Delaware jury found a CFAA violation over partly login-gated fare data in July 2024 and awarded exactly $5,000, the civil loss threshold; in January 2025 Judge William Bryson erased the verdict because only $2,457.72 of the claimed losses qualified as CFAA “loss” (Cooley, 2025). The appeal is pending at the Third Circuit as of July 2026. CFAA exposure concentrates behind logins and technical barriers, and even there the loss threshold is a litigated gate.

Copyright and fair use: content versus facts

Copyright protects scraped content, not scraped facts: Feist Publications v. Rural Telephone (Supreme Court, 1991) settled that facts and non-creative compilations carry no copyright. Fair use is a four-factor, case-by-case defense, not a safe harbor. The older trespass to chattels theory is still pleaded on occasion, but it requires proof of actual impairment to the server and has been mostly dormant in the modern cases.

The preemption point comes from X Corp v. Bright Data. Dismissing X’s claims in May 2024, Judge William Alsup wrote: “X Corp. wants it both ways: to keep its safe harbors yet exercise a copyright owner’s right to exclude, wresting fees from those who wish to extract and copy X users’ content” (CNBC, 2024). State-law claims over copying public data were held preempted by the Copyright Act. One caution: the parties settled confidentially in June 2025, so no final merits precedent exists; do not cite this case as “scrapers won.”

DMCA §1201: bypassing anti-bot measures is the new front

DMCA §1201 prohibits circumventing technical measures that control access to protected works. In October 2025, Reddit sued Perplexity AI, SerpApi, and Oxylabs in the Southern District of New York on this theory, over scraping Reddit content through Google’s search results (ZwillGen, 2026); the suit is pending as of July 2026. The practical read: bypassing rate limits, CAPTCHAs, or anti-bot systems now carries the highest theory-risk of any scraping behavior. robots.txt remains a norm rather than a law, but a documented pattern of ignoring it feeds these narratives in court.

Terms of Service: logged-in versus logged-off

Contract is the theory that actually ended hiQ, and it turns on whether you are bound: clickwrap terms bind account holders, browsewrap terms bind far less reliably. In Meta v. Bright Data (January 2024), Judge Edward Chen granted the scraping company summary judgment: “The Facebook and Instagram Terms do not bar logged-off scraping of public data; perforce it does not prohibit the sale of such public data” (Farella Braun + Martel, 2024). Non-account-holders, the court noted, do not see the terms and cannot be bound. Two qualifications: it is a district-court ruling tied to Meta’s specific terms, and Reddit v. Anthropic (filed June 2025, pending as of July 2026) is testing a broader contract-only theory. Terms exposure is civil, not criminal; as hiQ shows, civil is enough to end a data business. That is the federal half of the map. The other half is written state by state.

Quick Summary

Q: Which federal laws actually apply to web scraping in the US?

A: Four theories: the CFAA (narrow after Van Buren and hiQ for public, logged-off data), copyright (content, not facts, with preemption of state-law workarounds), DMCA §1201 (bypassing anti-bot measures, the emerging front), and terms-of-service contract claims (the theory that ended hiQ with a $500,000 judgment). Access method decides which theory bites.

Which States Regulate Scraped Personal Data in 2026?

Twenty states have comprehensive consumer privacy laws in effect as of July 2026; IAPP’s stricter count says 19, excluding Florida because its law applies only to companies with $1B+ global revenue. The 2026 wave alone brought three new laws live on January 1 (Indiana, Kentucky, Rhode Island), Connecticut’s SB 1295 amendments on July 1, and new California data-broker requirements on August 1. The map moves 2 to 3 times a year, hence the review stamp on the table.

A scraping team reads the table with three questions per state: does our dataset contain personal data of that state’s residents, do we trip the applicability threshold, and does the “publicly available” carve-out actually cover what we collected?

The 2026 map: 20 laws, one table

StateStatuteEffective dateCarve-out modelScraping-relevant notes
CaliforniaCCPA (amended by CPRA)Jan 1, 2020; CPRA amendments Jan 1, 2023CA-model (narrower)Revenue OR volume thresholds; biometric exclusion; broker registry; ADMT regs eff. Jan 1, 2026
VirginiaVCDPAJan 1, 2023VA-modelThe template most later states copied; opt-in consent for sensitive data
ColoradoCPAJul 1, 2023VA-modelRecognizes universal opt-out signals (GPC)
ConnecticutCTDPA + SB 1295Jul 1, 2023; amendments Jul 1, 2026VA-model35,000-consumer threshold; ANY sensitive-data processing or sale triggers scope
UtahUCPADec 31, 2023VA-modelBusiness-friendly variant; note the odd effective date
TexasTDPSAJul 1, 2024VA-modelNo volume threshold; data-broker registry; first AG enforcement (Allstate)
OregonOCPAJul 1, 2024VA-modelBroker registry; further provisions phased through Jan 1, 2026
FloridaFDBRJul 1, 2024VA-model$1B+ revenue gate; the reason IAPP counts 19
MontanaMCDPAOct 1, 2024VA-modelStandard VA-model mechanics
DelawareDPDPAJan 1, 2025VA-modelStandard VA-model mechanics
IowaICDPAJan 1, 2025VA-modelNarrower rights set than most peers
NebraskaNDPAJan 1, 2025VA-modelTDPSA-style structure
New HampshireNHDPAJan 1, 2025VA-modelStandard VA-model mechanics
New JerseyNJDPAJan 15, 2025VA-modelNote the mid-month effective date
TennesseeTIPAJul 1, 2025VA-modelStandard VA-model mechanics
MinnesotaMCDPAJul 31, 2025VA-modelStandard VA-model mechanics
MarylandMODPAOct 1, 2025VA-model text, strictest dutiesStrict-necessity minimization; flat ban on selling sensitive data
IndianaINCDPAJan 1, 2026VA-model2026 newcomer
KentuckyKCDPAJan 1, 2026VA-model2026 newcomer
Rhode IslandRIDTPPAJan 1, 2026VA-modelLow 35,000-consumer thresholds (MultiState, 2026)

Sources: MultiState (Feb 2026), IAPP (Jan 2026), and the IAPP US State Privacy Legislation Tracker. Last reviewed July 2026.

The 20 comprehensive state privacy laws in effect as of July 2026, grouped by the year they went live: California in 2020; Virginia, Colorado, Connecticut, and Utah by the end of 2023; Texas, Oregon, Florida, and Montana in 2024; Delaware, Iowa, Nebraska, New Hampshire, New Jersey, Tennessee, Minnesota, and Maryland in 2025; and the three January 1, 2026 newcomers Indiana, Kentucky, and Rhode Island highlighted. IAPP's stricter count is 19 because Florida's law applies only to companies with $1B+ global revenue, and Connecticut's SB 1295 amendments landed July 1, 2026. Sources: MultiState (Feb 2026), IAPP (Jan 2026).
The 20-state map, grouped by effective year

What “publicly available” actually exempts (narrower than scrapers assume)

Every one of the 20 laws excludes “publicly available information” from its definition of personal data, and the definitions differ. Keep two terms distinct: publicly accessible describes a page anyone can load; “publicly available” is a statutory term of art with conditions attached.

California’s is the narrower model. Under Cal. Civ. Code § 1798.140(v)(2), “publicly available” means: “(I) Information that is lawfully made available from federal, state, or local government records. (II) Information that a business has a reasonable basis to believe is lawfully made available to the general public by the consumer or from widely distributed media. (III) Information made available by a person to whom the consumer has disclosed the information if the consumer has not restricted the information to a specific audience.” The statute then subtracts: “‘Publicly available’ does not mean biometric information collected by a business about a consumer without the consumer’s knowledge.”

The Virginia-model states run broader but keep the same limit: content the consumer restricted to a specific audience does not qualify. Logged-in, followers-only, or privacy-limited content fails the carve-out in every state.

Public on the internet does not equal “publicly available” under privacy law. California expressly excludes biometric data collected without the consumer’s knowledge, and every state excludes content the consumer restricted to a specific audience.

Side-by-side comparison of the two statutory models for the 'publicly available' carve-out in state privacy laws. The California model is narrower: government records, widely distributed media, and consumer-disclosed information can qualify, but biometric information collected without the consumer's knowledge is excluded outright under Cal. Civ. Code section 1798.140(v)(2). The Virginia model, the template most later states copied, runs broader but keeps the same hard limit: content the consumer restricted to a specific audience never qualifies, so logged-in and followers-only data fails the carve-out in every state. Deletion and opt-out duties bind the data holder however the data was gathered. Source: Cal. Civ. Code section 1798.140(v)(2) and the state statutes cited in the article, as of July 2026.
Publicly accessible is not ‘publicly available’

One obligation attaches regardless of the carve-out analysis: deletion and opt-out duties bind the data holder however the data was gathered. Re-scraping a record a consumer already deleted has no clean statutory answer as of July 2026; treat it as an open operational question.

CCPA/CPRA deep dive: the California stack

The CCPA applies at $25M+ annual revenue, or personal information of 100,000+ consumers or households bought, sold, or shared, or 50%+ of revenue from selling or sharing personal information (IAPP, 2026). Scraped personal information triggers notice-at-collection duties, and “sale” is broader than cash: any transfer for “monetary or other valuable consideration” counts, a reading the California AG established against Sephora in 2022. California also recognizes the Global Privacy Control (GPC) opt-out signal, and its ADMT regulations took effect January 1, 2026.

Penalties run $2,500 per violation and $7,500 per intentional violation or violations involving minors’ data, and the CPPA’s Honda fine multiplied affected consumers by $2,500: per-consumer stacking is regulator practice, not a theoretical ceiling (Cal. Civ. Code § 1798.155; CPPA, 2025). The Delete Act’s broker mechanics get their own section below.

Texas TDPSA and the first-ever state enforcement (Texas v. Allstate/Arity)

The TDPSA has no volume threshold at all: any non-small-business processing Texans’ personal data is in scope, and even SBA-defined small businesses need consent to sell sensitive data (IAPP, 2026).

Then Texas used it. On January 13, 2025, the Texas Attorney General sued Allstate and its data subsidiary Arity, the first AG suit under any comprehensive state privacy law, alleging covert collection and sale of driving data covering 45M+ consumers and seeking up to $7,500 per violation under the TDPSA plus $10,000 per violation under the broker law. The case is pending as of July 2026: the filing and the penalties sought are the facts, no outcome exists. The signal is still legible: covert large-scale collection plus resale is now an attorney general priority.

The strict outliers: Maryland, Connecticut, and the minors’ rules

Maryland’s MODPA (effective October 1, 2025, applying to processing from April 1, 2026) rewrote the default: collection is allowed only as “reasonably necessary and proportionate” to a product or service “requested by the consumer,” a duty rather than a disclosure. A scraped dataset of Maryland residents was, by construction, not collected to provide something the consumer requested; commentators read MODPA as the state law most structurally hostile to bulk collection. It also carries the first flat ban on selling sensitive data; no consent cures it (EPIC, 2025).

Connecticut’s SB 1295 (effective July 1, 2026) cut the threshold from 100,000 to 35,000 consumers and extended the law to any entity processing sensitive data or selling personal data, regardless of volume (Covington, 2025). Even a small number of Connecticut residents’ sensitive records can now pull an out-of-state company into scope. Oregon’s minors’ amendments and Rhode Island’s low 35,000-consumer thresholds round out the tightening trend. And where the carve-outs subtract a category outright, the special-category statutes below take over.

Quick Summary

Q: Which US states regulate scraped personal data in 2026?

A: Twenty states have comprehensive privacy laws in effect as of July 2026 (19 by IAPP’s stricter count). All exempt “publicly available” data but define it narrowly: audience-restricted content never qualifies, and California excludes covertly collected biometrics outright. The strict outliers, Maryland’s necessity standard, Connecticut’s no-threshold sensitive-data trigger, and Texas’s no-volume-threshold law, are where scraped datasets most often land in scope.

Special-Category Data: Biometrics, Health, and Minors

Scraping liability concentrates by data category, not by page accessibility. Three categories have already produced nine-to-ten-figure outcomes or carry private rights of action.

Biometrics. Illinois’s BIPA (740 ILCS 14) requires informed written consent before capturing biometric identifiers, including faceprints computed from photos, with a private right of action at $1,000 per negligent and $5,000 per intentional or reckless violation (a 2024 amendment, SB 2979, made repeated scans of one person a single violation). Clearview AI is the priced example: a faceprint database built from reportedly 60+ billion public images ended in an Illinois class settlement, final-approved by Judge Sharon Johnson Coleman on March 20, 2025, valued at $51.75 million and paid as a 23% equity stake because the company lacked the cash; 22 state attorneys general and DC opposed it (Troutman, 2025). A 2022 ACLU consent decree separately bans Clearview from selling its faceprint database to most private businesses. Texas’s CUBI, AG-enforced at up to $25,000 per violation, produced Texas v. Meta: $1.4 billion on July 30, 2024, the largest single-state AG privacy settlement at the time (Texas AG, 2024). Washington’s HB 1493 (2017) is the third biometric statute. The screen: does the pipeline touch faces or photos convertible to faceprints?

Health-inferable data. Washington’s My Health My Data Act (in force March 31, 2024) regulates consumer health data outside HIPAA, broadly enough to reach inferences from searches, purchases, and location near clinics. It carries a private right of action via Washington’s Consumer Protection Act: actual damages plus treble damages up to $25,000. The first class action, Maxwell v. Amazon, was filed February 10, 2025 and is pending as of July 2026 (Orrick, 2025). Nevada’s SB 370 covers similar ground without the private right of action. The screen: “we don’t collect health data” fails if health status is inferable from what you do collect.

Minors’ data. Known-minor audiences carry sale and targeted-advertising bans across the newer state laws, with age-appropriate design codes spreading through 2026 amendments (Connecticut, Oregon). Adjacent regime, one line: names and likenesses in commercial products can trigger state rights of publicity independent of any privacy statute.

The through-line: public photos are not consent. Data category is one trigger; what a team does with the data downstream is the next.

Quick Summary

Q: Why are biometrics, health data, and minors’ data the highest-risk categories for scraping?

A: They carry consent requirements and private rights of action that public accessibility does not cure. Clearview AI’s faceprints from public photos settled at 23% of company equity ($51.75M value, 2025); Texas v. Meta settled biometric claims at $1.4B (2024); Washington’s My Health My Data Act lets consumers sue over inferred health data, with treble damages up to $25,000.

Data-Broker Registration: When Selling Scraped People-Data Makes You a Broker

Selling scraped personal data about consumers you have no direct relationship with presumptively makes you a data broker in the four registry states: California, Texas, Vermont, and Oregon. Registration is annual and separate from the privacy laws above.

California added teeth. The Delete Act (SB 362) built DROP, the CPPA’s one-stop deletion platform: registered brokers had to begin accessing it January 1, 2026 and must process all deletion requests by August 1, 2026, days after this article’s review date, at $200 per-deletion-request per-day penalties that compound across consumers (IAPP, 2026). It is the single most concrete compliance date in this article’s scope.

Registration failure is also the easy count for an enforcer: Texas charged Allstate with failure to register alongside its TDPSA counts, seeking $10,000 per violation on registration alone (Texas AG, 2025, pending). One caveat: “we only sell aggregated or derived data” does not automatically exit broker definitions; the statutory wording governs per state.

Quick Summary

Q: Do you have to register as a data broker if you sell scraped people-data?

A: In California, Texas, Vermont, and Oregon, yes, presumptively, if the consumers are ones you have no direct relationship with. California’s DROP deletion deadline of August 1, 2026 carries $200 per-request per-day penalties, and Texas already paired a $10,000-per-violation registration count with its first privacy enforcement suit.

The AI Layer: Training-Data Rulings and the 2026 AI Statutes

Reselling scraped data is one regulated downstream use. Training AI on it is the other. Two developments converged across 2025 and 2026: courts began ruling on fair use for AI training, and the first state AI statutes touching training data took effect January 1, 2026. The through-line of both is provenance.

The rulings. Bartz v. Anthropic split the question in two. In June 2025, Judge Alsup held that training LLMs on lawfully acquired books is fair use, but downloading from pirate libraries such as LibGen was not protected: the acquisition channel, not the training, created the liability. The settlement priced that channel: final approval on July 20, 2026 at $1.5 billion, the largest copyright class settlement in US history, roughly $3,000 per work across about 500,000 pirated books, with destruction of the pirated files ordered (Authors Alliance, 2026). Thomson Reuters v. Ross Intelligence points the other way: fair use rejected outright in February 2025, with the Third Circuit hearing argument June 11, 2026, the first federal appellate case squarely on the question. No decision exists as of July 2026. NYT v. OpenAI remains active with no merits ruling as of July 2026; its January 2026 order compelling production of 20 million de-identified ChatGPT logs shows courts treating AI pipelines as fully discoverable (Pillsbury, 2026).

“Fair use covers AI training” is not a safe assumption: one court blessed lawful copies, another rejected fair use outright, and the appellate answer is pending. The answerable question is narrower: can you document where every record came from? Teams sourcing corpora can start with our guide to compliant web scraping for AI training data.

The statutes. California’s AB 2013 requires generative AI developers to publish training-data documentation for models made available in California, effective January 1, 2026; scraped corpora are squarely in scope (King & Spalding, 2025). Texas’s TRAIGA took effect the same day, intent-based with a NIST AI RMF safe harbor. Colorado repealed its AI Act for the SB 26-189 ADMT law, signed May 14, 2026, effective January 1, 2027 (Miller Nash, 2026). Documentation duties are arriving faster than the fair-use question is resolving.

Promotional banner: training data with an acquisition trail per record. Courts priced pirated acquisition at $1.5B in Bartz v. Anthropic (final approval July 2026) and California's AB 2013 made training-data documentation law, so Forage AI builds data-for-AI pipelines around the acquisition trail, shown as a chain from source page through acquisition log to AI-ready dataset.
Data for AI with the acquisition trail attached

Quick Summary

Q: Is it legal to scrape data for AI training in the US?

A: Unsettled, and provenance-driven. Bartz v. Anthropic blessed training on lawfully acquired copies but priced pirated acquisition at $1.5B (final approval July 20, 2026); Thomson Reuters v. Ross rejected fair use, with the first appellate ruling pending as of July 2026; and California’s AB 2013 requires GenAI developers to publish training-data documentation from January 1, 2026. The defensible position is an acquisition trail for every record.

Enforcement: What Non-Compliance Has Actually Cost

Enforcement moved from theory to a priced reality across 2024 and 2025. The enforcers are state attorneys general, California’s CPPA, the FTC, and private classes where statutes allow them (BIPA, MHMD).

ActionEnforcerAmountDateWhy it matters for scraping teams
Sephora (CCPA)CA AG$1.2MAug 2022Historical first CCPA settlement; adtech data flows count as “sale”
DoorDash (CCPA/CalOPPA)CA AG$375KFeb 2024Marketing co-op data swap treated as a “sale” without notice
American Honda (CCPA)CPPA$632,500Mar 2025First major CPPA fine; consumers × $2,500 stacking math
Todd Snyder (CCPA)CPPA$345,178May 2025Broken opt-out mechanism; over-collection during rights requests
Healthline (CCPA)CA AG$1.55MJul 2025Largest CCPA settlement to date; article-reading data implied diagnoses
Texas v. Meta (CUBI)TX AG$1.4BJul 2024Biometrics without consent; largest single-state AG privacy settlement at the time
Texas v. Allstate/Arity (TDPSA + broker law)TX AGPending; seeks $7,500/violation + $10K/violationFiled Jan 2025First AG suit under a comprehensive state privacy law; covert bulk collection
Clearview AI (BIPA class)Private class$51.75M value (23% equity)Mar 2025Faceprints from public photos; equity-funded settlement
FTC: X-Mode & InMarketFTCConduct bansJan 2024Banned from selling precise/sensitive location data
FTC: Gravy Analytics & MobilewallaFTCConduct bansDec 2024Third-party data without verified consent treated as unfair practice

Sources: agency releases and law-firm alerts (CA AG 2022–2025; CPPA 2025; Texas AG 2024–2025; FTC 2024; Coblentz 2025; Cooley 2025). Table current as of July 2026.

Horizontal bar chart of verified privacy enforcement outcomes 2022 to 2025 on a log scale, grouped by enforcer type. Todd Snyder $345,178 (CPPA, May 2025); DoorDash $375,000 (California AG, February 2024); American Honda $632,500 (CPPA, March 2025); Sephora $1.2M (California AG, August 2022); Healthline $1.55M (California AG, July 2025, the largest CCPA settlement to date); Clearview AI $51.75M settlement value paid as a 23% equity stake (BIPA private class, final approval March 2025); Texas v. Meta $1.4B (Texas AG under CUBI, July 2024). Texas v. Allstate, filed January 2025, is pending with no outcome and is excluded. Sources: agency releases and law-firm alerts cited in the article (CA AG 2022-2025; CPPA 2025; Texas AG 2024; Troutman 2025), as of July 2026.
From $345K to $1.4B, on a log scale

Healthline’s $1.55M is the anchor for CCPA exposure: the largest settlement under the statute to date, over sharing article-reading data that implied medical diagnoses (CA AG, July 2025). The compounding mechanism sits in the Honda action: the CPPA computed the fine as affected consumers times $2,500. Small per-violation figures become large totals through ordinary regulator arithmetic.

The pattern across the dataset is consistent: sensitive and health-adjacent data (Healthline), biometrics (Texas v. Meta, Clearview), and covert bulk collection plus resale (Allstate, pending) draw the enforcement, not garden-variety public-page scraping. Two reading notes: Sephora is 2022 context, not current practice, and Allstate is a filing with penalties sought, not an outcome.

Quick Summary

Q: What has privacy non-compliance actually cost data collectors?

A: Verified 2024–2025 outcomes range from $345,178 (Todd Snyder, CPPA) to $1.4B (Texas v. Meta, biometrics), with Healthline’s $1.55M the largest CCPA settlement to date. Sensitive categories, biometrics, and covert bulk collection draw the enforcement, and per-violation stacking (Honda: consumers × $2,500) turns small statutory figures into large totals.

What Does Each Scraping Use Case Actually Trigger?

Same pipeline, different data, different legal surface. The map runs on three screening questions per project: does the extracted dataset contain personal data, and of which states’ residents; how is the source accessed (public logged-off, credentialed, or technically gated); and what is the downstream use (internal analytics, resale, AI training, commercial product)?

Risk ranks by data category, access method, and use. It does not rank by how public the page looks. That is the table’s whole argument.

Use caseLaws/theories implicatedRealistic risk postureRequired controls
Price, product, and market dataCopyright thin (facts, per Feist); CFAA low if logged-off; ToS if logged-inLowest-risk rowLogged-off access; no content republishing
Company/firmographic dataMostly non-personal; sole-proprietor and contact details edge into state lawsLow, with a personal-data edgeField minimization; threshold awareness
People data (contacts, profiles, lead lists)All 20 state laws; carve-out limits; “sale” breadth; broker registration if resoldHighest statutory surfaceRights-request and GPC handling; broker analysis; carve-out check per state
Images, media, and facesCopyright (content); BIPA/CUBI/HB 1493 if faceprints; publicity rights if commercialHighest severity per recordNo faceprint computation without consent posture; licensing review
AI training corporaCopyright/fair-use split (Bartz vs. Ross); AB 2013 documentation; provenance dutyUnsettled; provenance-drivenAcquisition-channel logging per record
Anything behind a login or paywallToS contract exposure (hiQ endgame); carve-outs fail; §1201 if barriers bypassedHighest behavioral riskCounsel sign-off before credentialed collection; no barrier circumvention

Sources: consolidated from the cases and statutes above (court records, state statutes, agency releases). Stamped as of July 2026.

Decision map walking three screening questions (data category, access method, downstream use) into six scraping use-case rows with their risk postures: price, product, and market data is the lowest-risk row (uncopyrightable facts, logged-off access); company and firmographic data is low risk with sole-proprietor and contact-detail edge cases; people data carries the highest statutory surface (all 20 state laws plus broker registration); images, media, and faces carry the highest severity per record under the biometric statutes; AI training corpora are unsettled and provenance-driven as of July 2026; and anything behind a login or paywall carries the highest behavioral risk through contract and DMCA section 1201 theories. Risk follows data category, access method, and downstream use, not how public the page looks. Source: the article's use-case map, consolidated from court records, state statutes, and agency releases, as of July 2026.
Risk follows the data, not the page

Row 1, price and market data, is the lowest-risk configuration: facts carry no copyright under Feist, and logged-off access keeps the CFAA and contract theories quiet. The posture holds only while it stays logged-off and republishes nothing.

Row 2, firmographic data, sits mostly outside the privacy statutes because companies are not consumers. The edge case: sole proprietors’ and named contacts’ details are personal data of the states they live in. Our firmographic data guide covers this category.

Row 3, people data, carries the most statutory surface: all 20 state laws, the narrow carve-outs, “sale” as any valuable consideration, and broker registration if resold. Threshold monitoring and rights handling stop being optional here.

Row 4, images and faces, is where severity per record peaks. A photo is content under copyright; a faceprint computed from it is a biometric identifier under BIPA, CUBI, and Washington’s HB 1493. Clearview’s 23%-equity settlement is this row read wrong (National Law Review, 2025).

Row 5, AI corpora, turns on provenance: Bartz priced pirated acquisition at $1.5B while blessing lawful copies, and AB 2013 requires the documentation regardless (Authors Alliance, 2026). The pending Third Circuit ruling in Ross could shift this row.

Row 6, logged-in or paywalled content, is where “public data” intuitions fail hardest: you accepted the terms, so contract exposure is live (the hiQ endgame); the carve-outs exclude audience-restricted content in every state; and bypassing barriers adds the §1201 theory. Judge Chen’s logged-off holding is the mirror image: the protection he described stops at the login wall (Farella Braun + Martel, 2024).

Quick Summary

Q: Which laws apply to each web scraping use case?

A: Price and product data implicates the least (uncopyrightable facts, logged-off access). People data implicates the most statutory surface (20 state laws, “sale” breadth, broker registration). Faces, AI corpora, and logged-in content each carry a dominant theory: biometric statutes, copyright plus provenance duties, and contract respectively. Risk follows data category, access method, and downstream use, not page publicness.

A Compliance Baseline for Scraping Teams and Scraped-Data Buyers

Everything above compresses into a program baseline with two halves: controls for teams that scrape, and diligence for teams that buy. Exposure does not transfer with the invoice.

Program controls for teams that scrape

The four gates, run before any collection starts:

  1. Robots/ToS posture: documented per source, on file, not in someone’s head.
  2. Login state: logged-off public-web collection only, unless counsel signs off on a credentialed exception.
  3. Technical barriers: no circumvention of anti-bot measures, rate limits, or CAPTCHAs.
  4. Data category: PII, biometric, health-inferable, and minors’ data flagged before collection, not after delivery.

Behind the gates sit the standing obligations: field-level minimization at extraction time, threshold monitoring by state, rights-request and GPC handling, provenance logging per record, and an annual law-tracking pass against the IAPP tracker (linked in the state table above), because the statutes move 2 to 3 times a year. Validation of what actually arrives belongs in the same loop; our data quality framework covers that layer.

One-page compliance baseline for web scraping teams: four gates run before any collection starts (robots and terms-of-service posture documented per source; logged-off collection only, with credentialed exceptions needing counsel sign-off; no circumvention of anti-bot measures, rate limits, or CAPTCHAs; and a data-category screen flagging PII, biometric, health-inferable, and minors' data before collection), plus five standing obligations (field-level minimization at extraction time, threshold monitoring state by state, rights-request and GPC opt-out handling, provenance logging per record, and an annual law-tracking pass because the statutes move 2 to 3 times a year). Source: the article's compliance baseline section, as of July 2026.
Four gates, five standing obligations

Vendor diligence for teams that buy scraped data

“Our vendor handles compliance” is not a control. Obligations attach to the data holder, and the FTC’s Mobilewalla and Gravy Analytics orders sanctioned exactly this posture: relying on third-party-sourced data without verified consent (FTC, 2024). What a buyer controls is the diligence conversation. Six questions for any scraping vendor:

  • Do you collect logged-off public data only?
  • How do you document provenance per record?
  • How do you handle deletion and opt-out requests downstream?
  • Which sources’ ToS positions have you assessed, and where are they documented?
  • What do you filter out at extraction time (fields, categories, minors’ signals)?
  • What contractual warranties do you offer on acquisition method?

These are the questions Forage AI invites from its own prospects as a managed extraction provider; a vendor that cannot answer them in writing has answered them anyway. As a managed service, Forage operates compliance-aware extraction practices along the lines this article maps: a logged-off public-web collection boundary, field-level minimization at extraction time, and provenance tracking per record. The contract follows from the questions: acquisition-method warranties, provenance documentation, deletion-passthrough terms. For what a managed provider should own end to end, see what managed web data extraction covers.

Promotional banner: the six vendor-diligence questions to ask any scraping vendor, which Forage AI invites as a managed extraction provider. Logged-off only? Provenance per record? Deletion passthrough? ToS positions on file? Extraction-time filters? Acquisition warranties? Exposure does not transfer with the invoice, and a vendor who cannot answer in writing has answered anyway.
Six questions for any scraping vendor

Quick Summary

Q: What does a defensible compliance baseline for web scraping look like?

A: Four gates before collection (documented robots/ToS posture, logged-off-only access, no barrier circumvention, a data-category screen) and standing obligations after it (field minimization, threshold monitoring by state, rights-request and GPC handling, provenance logging, annual law tracking). Buyers add a vendor-diligence question set and contract terms on acquisition method, because exposure does not transfer with the invoice.

Expert Insights

The strongest lines in the map above belong to the judges, statutes, and regulators; they are collected here.

Expert Insights

“X Corp. wants it both ways: to keep its safe harbors yet exercise a copyright owner’s right to exclude, wresting fees from those who wish to extract and copy X users’ content,” wrote Judge Alsup, adding that “free rein” over public web data “risks the possible creation of information monopolies that would disserve the public interest.” (Judge William Alsup, N.D. Cal., X Corp v. Bright Data, May 2024)

“If this is what passes as technological harm, then more CFAA absurdity is certain to come in the near future.” (Kieran McCarthy, founding partner, McCarthy Law Group, on Ryanair v. Booking; Technology & Marketing Law Blog, March 2025)

The statutory texts carry the argument. California’s carve-out has a built-in subtraction: “‘Publicly available’ does not mean biometric information collected by a business about a consumer without the consumer’s knowledge” (Cal. Civ. Code § 1798.140(v)(2)). Maryland inverts the model, permitting collection only when “reasonably necessary and proportionate” to a service “requested by the consumer” (MODPA, per EPIC, 2025). The two mark the edges of the 20-state range.

Three anchors price the heavy rows of the use-case map: Judge Chen’s holding that Meta’s terms “do not bar logged-off scraping of public data” bounds row 6 from the safe side (N.D. Cal., January 2024); Clearview’s $51.75M equity settlement prices row 4 (final approval March 2025); Bartz’s $1.5B settlement over pirated acquisition prices row 5 (final approval July 2026).

The FTC’s December 2024 orders against Mobilewalla and Gravy Analytics are the closest regulatory statement on scraped-data supply chains: acquiring location data from third parties without verifying consumer consent was treated as an unfair practice, and Mobilewalla’s order carried the first-ever prohibition on collecting data from real-time bidding streams (FTC, 2024). Buyers, not only collectors, sit inside the enforcement perimeter.

Frequently Asked Questions

Which US states have privacy laws that affect web scraping?

Twenty states have comprehensive privacy laws in effect as of July 2026; IAPP’s stricter count says 19, excluding Florida’s $1B+ revenue-gated law. The newest wave brought Indiana, Kentucky, and Rhode Island live on January 1, 2026, with Connecticut’s amendments following July 1. The full effective-date table sits in the state-layer section above.

Does the CFAA make scraping public data illegal?

No, not for public, logged-off data in the Ninth Circuit, under Van Buren’s gates-up-or-down test and hiQ. The catch: hiQ still lost on breach of contract and closed under a $500,000 consent judgment. CFAA claims persist for credentialed and gated access; the Ryanair appeal is pending at the Third Circuit as of July 2026.

Is publicly available data exempt from state privacy laws?

Only within each statute’s definition, which is narrower than the everyday meaning. Content the consumer restricted to a specific audience never qualifies, and California expressly excludes biometric data collected without the consumer’s knowledge. A page being loadable by anyone does not make its data “publicly available” in the statutory sense.

Can you scrape faces or photos in the US?

Photos as content raise copyright questions; converting them into faceprints triggers the biometric statutes (Illinois BIPA, Texas CUBI, Washington HB 1493), with consent requirements and, in Illinois, a private right of action. Clearview AI scraped only publicly posted images and its Illinois settlement was valued at 23% of the company’s equity. Public posting is not consent.

What penalties apply if a scraping program violates the CCPA?

The statute sets $2,500 per violation and $7,500 per intentional violation or violations involving minors’ data, and regulators stack these per consumer, as the Honda fine’s consumers-times-$2,500 arithmetic showed in 2025. Healthline’s $1.55M settlement (July 2025) is the current ceiling.

Do Terms of Service override the right to scrape public pages?

Terms bind those who accept them. Meta’s terms were held not to bar logged-off scraping of public data (Meta v. Bright Data, 2024), but logged-in scraping means you accepted the terms, and that contract exposure is how hiQ ended. Reddit v. Anthropic, pending as of July 2026, is testing how far a contract-only theory stretches.

Conclusion: The Map Is the Law

With no comprehensive federal statute, the operative US rulebook for scraped data is the map you just worked through: 20 state laws with their carve-outs and thresholds, four federal theories sorted by access method, and a use-case grid that ties them to the data your team actually collects. That was the promise at the top of this page, and it is now your working artifact: scope each project against the three screening questions, the four gates, and the state table.

Four pending matters can shift this page: Thomson Reuters v. Ross and Ryanair v. Booking at the Third Circuit, summary judgment in NYT v. OpenAI, and California’s DROP deadline on August 1, 2026. This page is reviewed against the IAPP tracker; last reviewed July 2026.

Teams that would rather have a managed provider own this operational surface can start with the vendor-diligence questions above, and with the build-versus-buy decision guide for the wider sourcing decision.

Related Articles

S
Written by
Sai Subramaniam
Data Infrastructure Enthusiast, Forage AI

Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.

Reviewed by the team of experts at Forage AI for accuracy and clarity.

Related Blogs

post-image

Compliance & Regulation in Data Extraction

July 24, 2026

US Web Scraping Laws in 2026: State Privacy Laws, Federal Law, and a Use-Case Map for Data Teams

Sai S

5 min read

post-image

AI Powered Solutions

July 24, 2026

RAG as a Service in 2026: Top 15 Platforms Compared

Sai S

5 min read

post-image

Web Data Extraction

July 24, 2026

Grepsr Alternatives: What Actually Fixes the Wall You Hit (2026)

Sai S

5 min read