Zillow Data Scraping: What You Can Collect and 5 Ways to Get It

Zillow Data Scraping: What You Can Collect and 5 Ways to Get It

The script pulled 40 listings on Tuesday. On Wednesday it returned a page asking you to press and hold a button to prove you are human, and every run since has done the same.

That is where most Zillow data scraping projects stall. It is also the wrong place to start. Zillow's apps and websites averaged 239 million monthly unique users in the second quarter of 2026, according to the company's own results, and the data behind those pages comes from several places, each with its own rules.

There are five routes to Zillow data. One is free, one is gated, and one is the scraper you were writing. This guide covers what data a Zillow page actually holds, what each route returns, where each one breaks, and when the data you need is not on Zillow at all.

Quick Digest
  • Three sources, one page: a Zillow listing combines MLS listing feeds, county tax records and Zillow's own Zestimate model, and each has its own lag and error.
  • Free official data exists: Zillow Research publishes market-level CSVs for free with attribution, but nothing property-level.
  • The official API is not a dataset: Zillow's API terms prohibit storing or caching the data, which rules it out for building a property database.
  • A working scraper can still be wrong: page structure drifts, listing feeds get cut, and off-market sales never appear, so measure coverage against an outside count.

What data you can get from Zillow

A Zillow property page, keyed by a Zillow property ID (ZPID), is the unit most people want. It looks like one record. In practice it is three records from three different systems sharing a layout.

Zillow publishes a Zestimate for more than 100 million US homes, which is why the page feels complete even for houses that have never been listed. The table maps each group of fields to its source.

Field group Examples Where it comes from How fresh it is Watch out
Listing facts Price, status, beds, baths, square footage, photos, description, listing agent and broker Multiple listing service (MLS) feeds via Internet Data Exchange (IDX), plus broker feeds Near real time while listed IDX feeds do not carry everything an agent sees in the MLS; fields can disappear by policy
Price history List, price-change and sale events Past listings plus recorded sales Lags the county recording Off-market sales show up only once recorded
Tax history Assessed value, annual tax County assessor records Assessed on each state's cycle Characteristics can be years old
Zestimate Value and Rent Zestimate Zillow's valuation model Recalculated by Zillow A model estimate with a published error rate, not a price
Location and context Coordinates, schools, neighbourhood Zillow and third-party data Varies Not a substitute for parcel boundaries
Identifiers ZPID, address Zillow Stable within Zillow ZPID does not join to county data; use address or parcel number

The three sources disagree for structural reasons, not because someone made a mistake. The assessment profession's own standard, from the International Association of Assessing Officers (IAAO), expects property characteristics to be reviewed "at least every 4 to 6 years" and states that "only physical reviews can correct data errors." So the square footage on the tax record and the square footage in the listing can both be correct for their own purpose and still differ. As one Sacramento appraiser, Ryan Lundquist, put it in 2020, assessor records are "generally reliable, but I'm just saying sometimes they're not."

The Zestimate carries its own error. Zillow's published nationwide median error is under 2% for homes on the market and around 7% for homes that are not, as of September 2026. Treat it as a model feature, not a sale price. Fields can also vanish by policy: offers of compensation left MLS systems on 17 August 2024 under the National Association of Realtors (NAR) settlement, so any series spanning that date has a seam in it. If you are new to the categories themselves, our overview of what real estate data covers is the place to start.

Diagram: a Zillow listing page draws on three sources. The listing feed (MLS and IDX) supplies price, status, beds, baths, photos and agent, near real time. County assessor and recorder records supply tax history, assessed value and recorded sales, with a lag of weeks to years. Zillow's model supplies the Zestimate and Rent Zestimate, a valuation estimate rather than a sale price. None of the three is complete.
One Zillow page, three sources, each with its own lag.

Does Zillow have an API?

Zillow does not offer a public API anymore. Zillow retired its legacy public API in 2021, and its free research database, ZTRAX, ended on 30 September 2023. What remains is one free route and one gated one.

Zillow Research downloads are the route most people skip. Zillow Research publishes CSVs covering home values (ZHVI), rents (ZORI), inventory, new listings, price cuts and days on market at metro, county and ZIP code levels, among others. Zillow describes the data as free for public use by consumers, media, analysts, academics and policymakers, with clear attribution to Zillow. ZHVI reflects the typical home in the 35th to 65th percentile of value, smoothed and seasonally adjusted. If your question is what prices and rents are doing in a set of markets, this answers it without touching a listing page.

Python: load a Zillow Research ZHVI file and read the last 12 months for one metro

import pandas as pd

# Metro-level ZHVI file downloaded from Zillow Research (CSV, one column per month)
zhvi = pd.read_csv("zhvi_metro.csv")

metro = zhvi[zhvi["RegionName"] == "Austin, TX"]
month_cols = [c for c in zhvi.columns if c[:4].isdigit()]
print(metro[month_cols[-12:]].T)  # last 12 months of typical home value

Official APIs through Bridge Interactive are the gated route. As of September 2026, Bridge is Zillow Group's data platform. MLS listing feeds on it require approval from each MLS, and Zillow's own datasets (public records, Zestimates, reviews) require Zillow to approve your stated use case. The detail that decides most projects sits in the terms: Zillow's API data terms say you may not store, cache or download the data except as the terms allow, and the public records licence covers calls made to serve end users immediately. That fits a live display app. It rules out building a stored property dataset for analysis.

For listing data outside Zillow, the underlying source is the MLS itself. There were 484 MLSs in the US at the end of 2025, and each grants access under its own licence. Our guide on how to evaluate a real estate data API covers that decision in depth.

How do you scrape Zillow?

Routes 1 and 2 are the official ones. The other three involve collecting the page yourself or paying someone to, and route 3, your own collector, is what most people mean when they ask how to scrape Zillow.

The five routes compared

Use the cheapest route that answers your question.

Route What you get Main limit Main cost driver Best when
1. Zillow Research downloads Market-level values, rents, inventory No property-level records Free You need trends, not homes
2. Official APIs (Bridge) Listings per approved MLS; Zillow datasets per approved use Approval per source; no storing data Approval time You are building a display app
3. Your own collector Whatever the page shows Blocking, drift, terms of use Engineering and maintenance You need a small, specific set and accept the risk
4. No-code tools and prebuilt scrapers or datasets A fixed field set, fast You inherit someone else's coverage, schema and legal posture Per-record or subscription fees You need a one-off pull
5. A managed data partner A custom dataset to your schema, across sources Contract and lead time Service fee Property data is part of your product

Steps for each route

Route 1: Zillow Research downloads

  1. Open the Zillow Research data page and pick the metric: ZHVI for values, ZORI for rents, or inventory and days on market.
  2. Choose the geography (metro, county or ZIP code) and download the CSV.
  3. Load it and filter to your regions and months, as in the code above.
  4. Credit Zillow wherever you publish the numbers.

Route 2: Official APIs through Bridge

  1. Write down the exact use case, because approval is judged against it.
  2. Apply through Bridge Interactive for the MLS feeds or Zillow datasets you need.
  3. Wait for each approval; every MLS decides on its own.
  4. Build the app to request data when it displays it, not to store it.

Route 3: Your own collector

  1. Read Zillow's terms of use and decide, with counsel if the data feeds a commercial product, whether the project fits them.
  2. Collect listing URLs and ZPIDs from search results for your area.
  3. Fetch each property page at a low, steady rate and save the raw HTML.
  4. Parse the __NEXT_DATA__ JSON into your fields, as shown in the next section.
  5. Flag missing fields and failed pages on every run, and re-check the parser when those counts move.

Route 4: No-code tools and prebuilt scrapers or datasets

  1. List the fields and geographies you need before you compare tools.
  2. Run a sample for one ZIP code and check it against the live pages.
  3. Check the refresh date, the field coverage and how the vendor collects the data.
  4. Scale up only once the sample matches.

Route 5: A managed data partner

  1. Write the schema: fields, geographies, refresh cadence and delivery format.
  2. List every source you want covered, not only Zillow: county records, MLS licences, permits.
  3. Agree on the quality checks and on what counts as a complete record.
  4. Review a pilot dataset before you commit to a recurring feed.

How a Zillow collector works

A collector has four stages: search results pages yield listing URLs and ZPIDs, each property page is fetched, its data is parsed, and the result is mapped to your schema and stored.

Zillow's pages are built on Next.js and carry their data as JSON inside a script tag with the id __NEXT_DATA__, which holds more than the rendered page: price history, tax history, schools and coordinates. Parse that JSON rather than CSS classes. As of September 2026 that is how the pages are built. It will not stay that way, so write the parser to fail loudly.

Python: parse a Zillow property page's embedded JSON and flag missing fields

import json
from bs4 import BeautifulSoup

REQUIRED = ["zpid", "price", "bedrooms", "bathrooms", "livingArea"]

def find_property(node):
    """Walk the embedded JSON and return the first object that looks like a property."""
    if isinstance(node, dict):
        if "zpid" in node and "price" in node:
            return node
        children = node.values()
    elif isinstance(node, list):
        children = node
    elif isinstance(node, str) and node.startswith("{"):
        try:
            return find_property(json.loads(node))  # some caches store JSON as a string
        except ValueError:
            return None
    else:
        return None
    for child in children:
        found = find_property(child)
        if found:
            return found
    return None

def parse_listing(html: str) -> dict:
    tag = BeautifulSoup(html, "html.parser").find("script", id="__NEXT_DATA__")
    if tag is None or not tag.string:
        return {"_status": "no_embedded_json"}   # a block page or a layout change
    prop = find_property(json.loads(tag.string))
    if prop is None:
        return {"_status": "no_property_object"}
    record = {f: prop.get(f) for f in REQUIRED}
    record["_missing"] = [f for f in REQUIRED if record[f] in (None, "")]
    record["_status"] = "ok" if not record["_missing"] else "partial"
    return record

The field names are the ones visible in public examples as of September 2026; confirm them against your own pages. The _status and _missing fields matter more than the parsing, because they tell you, run over run, that the page changed. Content that renders client-side needs a headless browser such as Playwright or Selenium. Store raw HTML so you can re-parse without re-fetching. We are not publishing code to get past bot detection, and Zillow's terms prohibit bypassing CAPTCHA and similar precautions.

Why Zillow blocks scrapers

As of September 2026, Zillow sits behind a commercial bot-defence product from HUMAN Security (formerly PerimeterX). Blocked sessions get a 403 or a "Press & Hold" challenge, which is why a script that works on a laptop often fails on a server. The defence scores the connection and the browser, not only the headers, so copying headers rarely fixes it. Our layer-by-layer guide to getting blocked covers the mechanics.

Two numbers put this in proportion. In a 2026 benchmark of 35 websites, end-to-end AI agents completed simple HTML pages every time but succeeded as little as 5% of the time behind CAPTCHAs. And an independent proxy review measured the same residential networks at over 99.7% against a small test endpoint and in the mid-70s against major retail, search and social websites in the same research cycle. A success rate is a property of the target, not of the tool.

Expert Insights

"If we struggle to get the code, we're spending most of our time bypassing an anti-bot and the advantage given by the parsing solution will be minimal, compared to the total time spent on the project."

Pierluigi Vinciguerra, The Web Scraping Club, 2024

That is the trade routes 4 and 5 make on your behalf. A prebuilt scraper moves the blocking and maintenance to a vendor you cannot see into. A managed partner moves it to a team that is contractually on the hook for the data, which is how Forage AI works: Forage AI delivers the data, not just the pipeline.

Forage AI promo: property data that holds up at scale. Custom pipelines across listings, county records and documents, checked by a 3x QA team before they reach your system. Talk to our expert.
Property data that holds up at scale.

What breaks after your Zillow scraper works

The block is the loud failure, and you know about it within minutes. Three quiet failures have a wider blast radius, because the job keeps reporting success while running in a degraded mode nobody sees.

Drift. Zillow changes its page structure without notice. The collector still returns 200s, your parser finds nothing at the old path, and fields come back empty. Row counts stay flat. The fix is to track field completion per run (a list of missing fields for every record) and alert when it drops below its recent baseline. Our piece on why extraction pipelines break quietly goes further.

Coverage. A scraper that captures every Zillow listing still has gaps.

Zillow is not the whole market

Zillow does not show every home for sale. Its listings depend on MLS feeds that can be cut or held back, and a share of homes sell without ever being listed.

Zillow's for-sale listings arrive through MLS and IDX feeds, and those feeds can be cut. On 20 May 2026 the Chicago-area MLS, MRED, cut its feed to Zillow and about 43,000 listings disappeared, until a federal judge ordered the feeds restored two days later. Other listings never reach the feed by design. NAR's delayed marketing exempt listing, in effect since 2025, keeps a listing out of IDX and syndication "for any period as allowed by the local MLS." Zillow's own Listing Access Standards exclude privately marketed homes that miss its rules. And a share of homes sell without being listed at all: one MLS's study of 2019 to 2020 sales in the Washington, Baltimore and Philadelphia metros found 26% (115,760 of 442,829) sold off-MLS. The measure that catches this is an outside count. Compare your listings per ZIP code against Zillow Research inventory or MLS statistics each month.

Lag. Speed costs coverage too. In a 2023 academic study measured against a platform's own database, collecting ten minutes after publication missed 2.1% of items and an eight-hour lag missed 5.4%. The items you miss are the short-lived ones, so the gap is not a random sample.

This is the maintenance a managed partner absorbs. Forage AI handles selector drift, anti-bot evolution, and schema changes as part of the service.

Quick Summary

Q: Is a scraper that runs without errors collecting all of Zillow's data?

A: No. Page drift empties fields while jobs stay green, listing feeds get cut or held back by policy, and fast-selling homes slip past slow collection. Only an outside count per ZIP code and a per-run field completeness check reveal the gaps.

Forage AI promo: your scraper says 200 OK, but is your data complete? A grid of records with six of 24 fields empty. Forage AI handles drift, anti-bot changes and schema updates as part of the service. Talk to our expert.
A green job is not a complete dataset.

This article is for informational purposes only and does not constitute legal advice. Consult a qualified attorney for legal guidance specific to your situation.

Whether you may collect Zillow data at all is a separate question from whether your collector works, and there is no settled answer. Start with the contract. Zillow's Terms of Use, updated 28 October 2025, prohibit "automated queries (including screen and database scraping, spiders, robots, crawlers, bypassing 'captcha' or similar precautions, or any other automated activity with the purpose of obtaining information from the Services)."

The case most often cited to say public data is fair game is hiQ v. LinkedIn. In 2022 the Ninth Circuit wrote that "without authorization" does not apply to public websites, but that was on a preliminary injunction, and the case ended in December 2022 with a $500,000 consent judgment against hiQ, a permanent injunction and an order to destroy data derived from the scraped profiles. The Supreme Court's Van Buren decision left open whether contract terms can make access unauthorised. In Ryanair v. Booking.com, a jury found a Computer Fraud and Abuse Act (CFAA) violation for screen scraping in 2024, and the judge set the verdict aside in 2025 because the loss threshold was not met; the appeal is pending as of September 2026.

Kieran McCarthy of McCarthy Legal Group, who litigates scraping cases, wrote after the verdict was set aside in 2025: "If this is what passes as technological harm, then more CFAA absurdity is certain to come in the near future."

Photos are a separate and clearer question. In VHT v. Zillow, the Ninth Circuit held Zillow itself liable in 2019 for how it catalogued and displayed a photographer's listing images, and statutory damages were affirmed in 2023. In July 2025 CoStar sued Zillow over roughly 47,000 photos, raised to about 53,000 in a March 2026 filing; Zillow disputes the claims and the case is unresolved. If Zillow can be sued over listing photos, a scraper that copies them carries the same exposure.

Public is not the same as permitted, and following robots.txt is not a safe harbour. Agent names and phone numbers are personal data. Take four questions to counsel: how you access the pages, what the terms say, what content you copy, and what you do with it. Our overview of US web scraping laws covers the wider case law.

When the data you need is not on Zillow

Legal exposure is one reason teams look past the portal. Coverage is the other. Zillow is a consumer portal, and the deeper or harder the question, the less of the answer lives there.

  • Ownership and transfers live with the county recorder. Deeds and liens can take two weeks or longer to become searchable after they are filed, so the newest transfers are missing from every downstream source for a while.
  • Assessor records sit with roughly 3,143 counties and county equivalents, each publishing on its own terms. One open project tracking 1,971 US address sources merged 183 fixes for broken sources in about 13 months.
  • Permits and title documents are often PDFs that need document processing, not scraping.
  • Commercial property data sits largely outside consumer portals, with the largest provider reporting $3.2 billion in revenue in 2025.
  • Off-market sales show up only in recorded deeds.

These sources join on parcel number and address, not ZPID, and they break without notice. This is the work Forage AI takes on for teams whose product depends on property data: custom pipelines across county, listing and document sources, normalised to your schema. Every Forage AI delivery passes a 3x QA team before it lands in your system, and the first dataset typically arrives within 1-2 weeks of sign-off. For what to do with the data once it lands, see how teams automate property market intelligence, or how real estate investors use web data.

Forage AI promo: hard-to-find property data, delivered clean. County deeds, assessor rolls, permits and title documents, normalised to your schema, with the first dataset in 1-2 weeks. Talk to our expert.
Hard-to-find property data, delivered clean.

Keeping a property dataset honest

Pick the route whose limits you can live with, then watch those limits. For most teams that means three habits: an outside count per ZIP code every month, a field completeness check on every run, and one named person who watches for MLS and policy changes. None of it is glamorous. All of it is cheaper than finding the gap in a customer's report.

If you run a Zillow collector today, check one ZIP code this week against Zillow Research inventory. The difference is your coverage, and it is the number worth sharing with your team.

Frequently asked questions

Does Zillow have a public API?

No. Zillow retired its public API in 2021. Official access now runs through Bridge Interactive, where MLS feeds need each MLS's approval and Zillow's own datasets need approval of your use case. Zillow's API terms also prohibit storing the data.

Can I download Zillow data for free?

Yes, at market level. Zillow Research publishes free CSVs of home values, rents, inventory and days on market at metro, county and ZIP code levels, with attribution. It does not publish property-level records.

Why does Zillow keep blocking my scraper?

Zillow uses commercial bot defence that scores your connection and browser, not only your headers. A script that works locally often gets a 403 or a "Press & Hold" challenge from a server. Zillow's terms also prohibit automated access.

Can I use photos I scrape from Zillow?

Treat them as copyrighted. Listing photos belong to photographers, brokers or other rights holders, and Zillow itself was held liable for statutory damages in VHT v. Zillow over how it used them. Get rights from the owner.

How accurate is Zillow's Zestimate?

By Zillow's own published figures as of September 2026, the nationwide median error is under 2% for homes on the market and around 7% for homes that are not. Half of estimates are further off than that.

Does Zillow have every home for sale?

No. Its listings depend on MLS and IDX feeds that can be cut or held back, delayed-marketing listings are withheld by design, and a share of homes sell without being listed.

S
Written by
Sai Subramaniam
Data Infrastructure Enthusiast, Forage AI

Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.

Reviewed by the team of experts at Forage AI for accuracy and clarity.