Data Extraction

Top 10 Legal Data Providers: How Intelligence Teams Buy Legal Data (2026)

August 21, 2026

5 min read


Top 10 Legal Data Providers: How Intelligence Teams Buy Legal Data (2026) featured image

Buy the wrong class of legal data and no coverage number on the contract will save you. The recurring mis-buy: a market-intelligence team needs docket outcomes and signs with a firmographic directory, because both call themselves legal data providers. The category name hides four vendor classes: court-record and docket aggregators, firm and attorney directories, docket and outcome analytics, and custom extraction and managed feeds. Of the 17 pages ranking for this keyword on 19 August 2026, 13 mix those classes into one numbered list, and one marketplace’s “attorney data” page returns property, mortgage and aviation vendors at slots 3, 4, 6, 7 and 8.

So this article routes you to a class before naming a single logo, then gives you a coverage test runnable on a vendor sample in a week. By the end, you can name your class, the gap that class is structurally guaranteed to have, and what you will request in a vendor sample and measure on it.

Quick Digest

  • The four classes: docket aggregation, firm and attorney intelligence, outcome analytics, and custom extraction. Route your question to a class first.
  • The sources: federal filings come from one system, PACER; state filings from 50-plus systems tracked through 200+ reporting units.
  • The list: Forage AI leads for uncatalogued jurisdictions; UniCourt, Trellis, Docket Alarm and the Free Law Project cover dockets; Lex Machina, Westlaw and Bloomberg Law cover analytics; ALM and Leopard Solutions cover firm intelligence.
  • The test: a five-field sample request plus five measurements, runnable in a week, benchmarked against published court data.
  • The gaps: every class has a permanent blind spot, sourced to the judiciary, a federal rule, or the vendor’s own documentation.
  • The money: five pricing shapes, one published price anchor, and the licence clause that kills derivative products.
  • The routing: a job-to-class lookup, because “which legal database is best” is the wrong question.
2026 Edition · Strategic Guide
How to Get Started With Your Data Acquisition Strategy For AI
A strategic guide for data leaders who don’t know where to start.
Most guides about data infrastructure jump to the technical fix. This one starts a step earlier, at the strategy decision. It helps you see where you stand on the data acquisition maturity curve, what your options are, and what to ask before you pick a partner.
5 Data Acquisition Stages
3 Data Solutions
15 Min Read
Download the e-book
Free. Sent straight to your inbox.
We’ll email you the guide. No spam, unsubscribe anytime.
Two-by-two map of the four legal data vendor classes: docket aggregators, firm directories, outcome analytics and custom extraction, each with its routing question, what it emits, its structural blind spot and the providers in that class.

Four vendor types share one category name and emit different records: a docket feed emits filings, parties and dispositions; a directory, people and firms; an analytics product, derived metrics; custom extraction, whatever schema you scoped. The classes describe what a product emits, not what a company is, so several vendors carry two class tags here. Everything in this article is US-scoped.

The routing test takes one line: write your question as a sentence, then ask which class emits the record that answers it. “What was filed and when” is Class 1; “who are the firms and who moved,” Class 2; “what happened and how often,” Class 3; “nobody catalogues my jurisdiction,” Class 4. A bigger catalogue does not fix a class mismatch.

One scale marker, with a wrinkle: the American Bar Association‘s National Lawyer Population Survey counts 1,322,649 resident active attorneys for data year 2024; the ABA’s own Profile of the Legal Profession (December 2025) reports 1.35 million for the same year. Two publications from one body, roughly 27,000 apart: this article’s thesis in miniature. We use the NLPS figure.

1,322,649 resident active US attorneys, data year 2024. Source: American Bar Association, National Lawyer Population Survey.

The one-line test: which class does your question belong to?

The details the table compresses: coverage of an index is not access to documents (Class 1); refresh cadence and entity resolution decide Class 2; ask for the per-judge n before you believe a Class 3 win rate. Class 4 builds legal data extraction to your schema, and is the wrong purchase when a catalogue already covers your jurisdiction well.

Class What it emits The question it answers Structural blind spot Typical delivery
1. Court-record and docket aggregators Dockets, filings, parties, dispositions What was filed, when, by whom? Sealed, paper-only, and document tiers behind the index API, bulk, alerts
2. Firm, attorney and legal-entity directories Attorney and firm profiles, headcount, movement Who are the firms, and who moved? Whatever the licensing bodies did not report Platform, feed
3. Docket, outcome and contract analytics Win rates, time-to-rule, judge analytics What happened, and how often? Every gap in its source layer, plus small-n venues Platform, gated API
4. Custom extraction and managed feeds Whatever schema you scope Nobody catalogues my jurisdiction No catalogue to browse; scoping precedes data Managed feed

One standard to hold onto: the SALI Alliance‘s Legal Matter Standard Specification (LMSS), an 18,000-plus-tag MIT-licensed open ontology, rolling out v3 (referenced March 2026). It is the join key when you merge two vendors’ feeds.

Quick Summary

Q: What are the four kinds of legal data provider, and how do I tell which one I need?

A: Legal data providers fall into four classes: court-record and docket aggregators, firm and attorney directories, docket and outcome analytics, and custom extraction and managed feeds. Write your question as a sentence, then ask which class emits the record that answers it. The classes describe what a product emits, not what a company is, so several vendors sit in two of them.

Every vendor here buys, scrapes or is handed the same upstream records, so the upstream layer decides what any of them can sell you. The structural fact: one federal system with one electronic docket, against 50-plus independent state systems that never agreed on formats or definitions. That asymmetry is permanent.

Two-pane comparison of legal data sources: one federal system, PACER, at ten cents per page with a three-dollar document cap, versus more than fifty independent state systems tracked through over 200 reporting units.

PACER, the PCL API, and what it will not give you

Federal access runs through PACER at $0.10 per page, capped at $3.00 per document, with no fee owed below $30.00 in a quarter (effective January 2020). The cap does not apply to name searches, so a broad party-name sweep can cost more than pulling the documents it finds.

$0.10 per page, capped at $3.00 per document. The PACER fee schedule in force, effective January 2020; the cap does not apply to name searches. Source: Administrative Office of the U.S. Courts.

Does PACER have an API? Yes, and it is an index, not a document store: the PACER Case Locator (PCL) API, per its November 2024 user guide, returns immediate searches in groups of 54 with a 5,400-item ceiling, caps batches at 108,000 items, and bills every production search. Here is one party hit; note the searchFee field, and the absence of any document:

{
  "content": [
    {
      "courtId": "ilndc",
      "caseId": 306781,
      "caseYear": 2015,
      "caseNumber": 1445,
      "lastName": "Henderson",
      "firstName": "Nicholas",
      "partyType": "pty",
      "partyRole": "dft",
      "jurisdictionType": "Civil",
      "courtCase": {
        "courtId": "ilndc",
        "caseType": "cv",
        "caseTitle": "Lytx, Inc. v. Sanderson",
        "dateFiled": "2015-02-17",
        "effectiveDateClosed": "2015-03-12",
        "natureOfSuit": "890",
        "caseNumberFull": "1:2015cv01445"
      }
    }
  ],
  "pageInfo": { "number": 0, "size": 54, "totalPages": 1, "totalElements": 2 },
  "receipt": { "searchFee": ".10" }
}

Fifty-plus state systems, and why coverage is ragged

There is no nationwide database of court cases, and the states are the reason. The National Center for State Courts‘ Court Statistics Project assembled its 2024 dashboards (November 2025) from over 200 reporting units. Its reporting guide, revised 3 March 2026 and now 105 pages of counting rules, records “profound differences in how states defined and reported their caseload data.” And that is the tidier statistics layer; the record layer a vendor must gather carries its own legal considerations for scraping public records.

Hold the asymmetry as a ratio.

Even the market leader’s state coverage is about a hundred courts against complete 94-of-94 federal district coverage and 200-plus state reporting units. State-heavy work needs a specialist or a multi-vendor merge layer, not a bigger catalogue.

Open corpora: RECAP, CourtListener, and what “free” costs you

For federal research, the closest thing to a free alternative to PACER is the Free Law Project‘s CourtListener and RECAP archive, holding nearly every federal case via 30,000-plus contributors; the caveats (billed retrieval, a rate-capped federal-only Fetch API, a no-derivatives licence) are unpacked in entry #5 and the pricing section.

Quick Summary

Q: Where do legal data providers actually get their data?

A: Federal filings come from one system, PACER and CM/ECF, at $0.10 per page with a $3.00 per-document cap that does not apply to searches. State filings come from more than fifty independent systems whose own national statistics project assembles its picture from over 200 separate reporting units. The official PACER Case Locator API is an index of cases and parties, not a document store, and every production search on it is billable.

How we ranked these ten

Every vendor below resells some slice of that upstream layer. Five criteria decided the top legal data providers below. One: does it sell data, or software that consumes data? That test removes roughly 40% of the entries on competing lists: contract-lifecycle, e-discovery, practice-management and legal-AI products, disqualified as a group. Two: is the coverage claim verifiable on a page we could open? Every figure is printed as published, attributed, retrieved 19 August 2026, and labelled a vendor claim. Three: can a data team consume it? Four: is the pricing shape and licensing posture knowable? “Publishes no list price” is a finding, not a gap. Five: is the entry consolidation-accurate as of August 2026? Acquired vendors merge into acquirers, one live rebrand is corrected, and defunct products are excluded: the inventor of state-court analytics reached 25 states, raised $5.7M and shut down in June 2022, yet two live directories still list it.

Three disclosures. Patent and trademark data is scoped out: a different buying motion. Two coverage claims could not be verified because the vendors’ marketing pages are bot-gated; we say so rather than reprint. And Forage AI publishes this article and holds the #1 slot in the managed-extraction category: a class placement, not a quality leaderboard, no vendor paid for placement, and the other nine entries were written from their own published documentation.

Numbered list of the five ranking criteria used to select the top 10 legal data providers: sells data rather than software, verifiable coverage claims, consumable delivery, knowable pricing shape, and consolidation accuracy as of August 2026.

Legal data providers at a glance

# Provider Class Best for
1 Forage AI 4 (custom extraction and managed feeds) The jurisdiction, registry or portal no catalogue covers, delivered as a maintained feed
2 UniCourt 1 (docket aggregator) API-first federal and state docket normalisation for teams building on the feed
3 Trellis Research 1 (state trial specialist) State trial court work, where the big platforms structurally are not
4 Docket Alarm (vLex / Clio) 1 + 3 Mid-market docket tracking and alerting with a real published price
5 Free Law Project (CourtListener & RECAP) 1 (open corpus) Federal research and prototyping where you can live with the licence
6 Lex Machina (LexisNexis) 3 (outcome analytics) Federal litigation outcome analytics with complete district coverage
7 Thomson Reuters / Westlaw 3 + 2 Research plus derived analytics in one relationship
8 Bloomberg Law 3 + 2 Enterprise research with a docket API attached, if you can absorb an opaque price
9 ALM Intelligence / Law.com Compass 2 (firm intelligence) Law-firm financials, client representation and market structure
10 Leopard Solutions (SurePoint Legal Insights) 2 (firm intelligence) Attorney movement and firm headcount, refreshed twice weekly

Category A. Class 4: custom extraction and managed feeds

This category exists because of what the rest of the article documents: document tiers that stop at the county line, commonwealth courts outside every framework, portals nobody indexes. Two other vendors on this SERP, Grepsr and APISCRAPY, sell general-purpose managed web-data extraction; neither publishes a legal-vertical coverage record we could verify on their own pages (19 August 2026), and neither takes a numbered slot.

1. Forage AI

Attribute Detail
Vendor classClass 4: custom extraction and managed feeds
Coverage (as published, retrieved 19 Aug 2026)Not a catalogue: coverage is scoped per engagement. Company operating scale: 500M+ websites crawled, 10M+ documents parsed.
DeliveryManaged feed, delivered to your schema (API, bulk, or scheduled delivery)
Refresh cadenceSet per scope in the engagement, not a fixed catalogue cadence
Pricing signalPublishes no list price; priced per scope, not per seat
Licensing & redistributionClient owns all extracted data; data is never resold or reused across clients
Best forThe jurisdiction, county portal or registry no catalogue covers, delivered as a maintained feed rather than a scraping tool
Not forTeams whose matters sit entirely inside an existing catalogue’s coverage; buying Class 4 there pays for scoping you do not need
What to ask for in the sampleA scoped pilot on one jurisdiction you choose, measured with the five-field test below
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)No rating with a published sample size located. Capterra listing exists with 0 reviews; Trustpilot shows roughly 1 review with no verifiable rating. (G2’s “Forage” listing at 4.5/5, n=18, is theforage.com, a different company, and is not usable.)
Evidence qualityNo self-serve review base; as a managed service, evaluation happens through scoped pilots rather than review platforms
What reviewers praiseNot printable from review platforms
What reviewers complain aboutNot printable from review platforms
SourceCapterra, Trustpilot (checked 19 Aug 2026)

What it is. Forage AI is a managed data extraction provider that owns source discovery, extraction, normalisation, QA and maintenance end to end: it delivers the data, not just the pipeline. The fit for legal’s uncatalogued end: custom pipelines built to your schema, a 3x QA team on every delivery, 1-2 weeks from sign-off to first dataset, and selector drift, anti-bot evolution and schema changes handled as part of the service.

Best for. The buyer whose jurisdiction failed the catalogue check, the exact gaps the category intro names. Not for all-federal dockets; buy Class 1 and stop reading this category.

What customers say. There is no published review base to report, and we will not invent one: Capterra shows zero reviews, Trustpilot roughly one, no verifiable rating (19 August 2026). Typical for a managed service evaluated through scoped pilots, and exactly why the sample test matters more here than a star rating.

Forage AI promotional banner: when no catalogue covers your jurisdiction, scope a managed extraction feed built to your schema, with QA on every delivery.

Category B. Class 1: court-record and docket aggregators

The filing layer: four routes to what courts emit, differing on state depth, delivery, licence and published price.

2. UniCourt

Attribute Detail
Vendor classClass 1: court-record and docket aggregator
Coverage (as published, retrieved 19 Aug 2026)Not reprinted here: unicourt.com’s coverage pages are CAPTCHA-gated to automated verification, so its published totals could not be confirmed on the vendor’s own pages
DeliveryAPI-first (Enterprise API, documented publicly at docs.unicourt.com)
Refresh cadenceNot published
Pricing signalPublishes no list price; limits and volumes set in the licence agreement
Licensing & redistributionContract terms; monthly API ceiling “agreed upon in your license agreement” per its own docs
Best forAPI-first federal and state docket normalisation for a team building on top of the feed
Not forTeams that need to size the integration from public documentation before talking to sales
What to ask for in the sampleYour actual contractual limits in writing: monthly ceiling, per-endpoint concurrency, case-tracking cap, daily document-order cap
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)Capterra (UniCourt Enterprise API) 3.0/5 (n=2); G2 listing exists with 0 reviews
Evidence qualityOne platform with a rating, and n=2 is too small to mean much; treat as anecdote
What reviewers praiseConsolidates court portals into one searchable interface; API suits direct integration and lead-gen search
What reviewers complain aboutOne reviewer felt misled about what case access the subscription included; consumer-side Trustpilot complaints about record-removal handling (directional only, no verifiable rating)
SourceCapterra, G2 (attributed to the platforms, not the vendor)

What it is. UniCourt is the most-named pure court-data vendor on this SERP: an API-first aggregator normalising federal and state dockets; its unusually candid technical docs, not its marketing, carry this entry.

Best for. A data team building on a normalised feed. The watch-out is contractual opacity: the limits that size your integration are not published anywhere.

What customers say. The printable review base is thin: Capterra 3.0/5 (n=2), G2 empty. The one substantive complaint, a reviewer who felt misled about included case access, teaches the test’s lesson: get scope in writing.

3. Trellis Research

Attribute Detail
Vendor classClass 1: state trial court specialist
Coverage (as published, retrieved 19 Aug 2026)State trial court coverage across 45 states, per the vendor’s own support knowledge base; larger figures on its marketing pages could not be verified (Cloudflare-gated) and are not reprinted
DeliveryWeb platform; delivery mechanics beyond the UI could not be verified on the gated main website (19 Aug 2026)
Refresh cadenceNot published
Pricing signalPublishes no list price
Licensing & redistributionNot stated publicly
Best forState trial court work, where the big platforms structurally are not
Not forTeams that need document retrieval outside the enumerated counties, or bulk delivery guarantees
What to ask for in the sampleThe document tier (free / requestable / purchasable) for every county your matters touch, in writing
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)Capterra 5.0/5 (n=2). (G2’s “Trellis” listing is an ecommerce advertising tool, a different company, and is not usable.)
Evidence qualityOne platform, n=2; a perfect score on two reviews is a sample-size artefact, not a verdict
What reviewers praiseEasy navigation with quick, thorough search; responsive support and leadership
What reviewers complain aboutNeither reviewer cited a complaint, but n=2 cannot support a “no complaints” claim
SourceCapterra (attributed to the platform, not the vendor)

What it is. Trellis Research is the state trial court specialist. Its own support knowledge base documents 45-state coverage, with documents tiered as free, free-to-request, and purchasable.

Best for. State-heavy dockets, because nobody else makes state trial courts the core product. The watch-out is the document tier: requestable documents exist only in Los Angeles and Cook counties. The index is wide; the paper is county-by-county.

What customers say. Capterra shows 5.0/5 (n=2); hold it lightly and test the counties instead.

4. Docket Alarm (vLex / Clio)

Attribute Detail
Vendor classClass 1 + 3: docket tracking with analytics
Coverage (as published, retrieved 19 Aug 2026)Federal dockets plus state coverage; reviewers name state and county gaps (New York and Massachusetts called out on Capterra)
DeliveryWeb platform with search, alerts and tracking; PACER documents passed through at cost
Refresh cadenceNot published
Pricing signal$99 per user per month flat fee; $39.99/mo pay-as-you-go plus $4 per document; enterprise requires teams of five or more (as of August 2026)
Licensing & redistributionNot stated publicly
Best forMid-market docket tracking and alerting with a real published price
Not forBulk data supply for product builds; the price is for the platform, not the feed
What to ask for in the sampleThe coverage list for the specific state and county courts your alerts depend on
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)Capterra 4.5/5 (n=13); G2 listing exists with 0 reviews
Evidence qualityOne platform with a usable sample; n=13 is the largest verified-on-page sample among the pure docket vendors here
What reviewers praiseFar more user-friendly than PACER, with keyword search across dockets; automated case alerts that run themselves
What reviewers complain aboutGaps in state and county court coverage (NY, MA named); per-document costs add up, including one report of repeat charges for the same document
SourceCapterra, G2 (attributed to the platforms, not the vendor)

What it is. Docket Alarm, now inside the Clio stack alongside vLex and Fastcase, owns the one genuinely published price in this category: $99 per user per month flat, with PACER fees billed separately at cost, plus a $39.99 monthly pay-as-you-go tier at $4 per document.

Best for. Tracking and alerting at a knowable cost. The published price is the platform, not the acquisition cost: PACER rides on top.

What customers say. Capterra shows 4.5/5 (n=13), the strongest verified base among the pure docket vendors. Praise: usability against raw PACER, and alerts. Complaints: state gaps and accumulating per-document charges.

5. Free Law Project (CourtListener & RECAP)

Attribute Detail
Vendor classClass 1: open corpus (nonprofit)
Coverage (as published, retrieved 19 Aug 2026)9M+ decisions from 2,000+ courts, stated as more than 99% of US precedential case law; RECAP archive of nearly every federal case, contributed by 30,000+ people; 16,000+ judges
DeliveryREST API v4 and bulk data; PACER Fetch API capped at 30 requests per minute, federal-only
Refresh cadenceContinuous, contributor-driven
Pricing signalThe archive is free; retrieval through PACER Fetch is billed by the court against your own PACER credentials
Licensing & redistributionCreative Commons BY-ND 4.0: attribution, no derivatives
Best forFederal research and prototyping where you can live with the licence
Not forProducts that redistribute derivatives; state trial court work; anything requiring sealed-item awareness
What to ask for in the sampleNothing to request: pull 20 of your own matters and count the hits yourself
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)No rating with a published sample size located. Capterra listing exists with 0 reviews; no G2, Trustpilot or TrustRadius listing found
Evidence qualityNo review base, as expected for a nonprofit open-source project; credibility signals live elsewhere (research guides, its open GitHub community)
What reviewers praiseNot printable from review platforms
What reviewers complain aboutNot printable from review platforms
SourceCapterra (checked 19 Aug 2026)

What it is. The Free Law Project publishes the largest open legal corpus in the country: over nine million decisions from more than 2,000 courts, the RECAP federal docket archive, and a judges database, behind a documented REST API. Requesting a federal docket looks like this; note your own PACER credentials are the ones billed:

curl -X POST \
  --data 'request_type=1' \
  --data 'pacer_username=xxx' \
  --data 'pacer_password=yyy' \
  --data 'docket_number=5:16-cv-00432' \
  --data 'court=okwd' \
  --header 'Authorization: Token <your-token-here>' \
  "https://www.courtlistener.com/api/rest/v4/recap-fetch/"

Best for. Federal research and prototyping, with three caveats: 30 requests per minute, federal-only; the corpus cannot distinguish sealed filings and its documentation says not to request them; and the licence is CC BY-ND 4.0, no derivatives. Not for anything that redistributes derived data.

What customers say. No published review base exists, unsurprising for a nonprofit. A 30,000-contributor archive and an open GitHub community are a different kind of evidence than a star rating.

Category C. Class 3: docket, outcome and contract analytics

The derived layer: all three answer “what happened and how often,” inherit their source layer’s gaps, and none is a bulk data supply.

6. Lex Machina (LexisNexis)

Attribute Detail
Vendor classClass 3: litigation outcome analytics
Coverage (as published, retrieved 19 Aug 2026)All 94 federal district courts, 13 courts of appeal, PTAB and specialty venues; over ten million cases; 45M customer-facing documents (vendor figures, stated as of April 2025)
DeliveryAnalytics platform; API access is enterprise-gated, not price-gated
Refresh cadenceNot published
Pricing signalPublishes no list price
Licensing & redistributionAnalytics subscription; underlying records are not a bulk export product
Best forFederal litigation outcome analytics with complete district coverage
Not forState-heavy work: enhanced state coverage reached roughly 100 courts (milestone announced May 2024) against 200+ state reporting units
What to ask for in the sampleThe per-judge sample size (n) behind any win rate in your venues
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)Capterra 2.0/5 (n=1); no usable G2 rating surfaced
Evidence qualityA 2.0/5 on n=1 is statistically meaningless; the honest reading is a minimal review-platform footprint for an enterprise-sold product
What reviewers praiseNot printable at this sample size
What reviewers complain aboutThe single reviewer had reservations about overall value and support (n=1)
SourceCapterra (attributed to the platform, not the vendor)

What it is. Lex Machina defined litigation analytics: win rates, time-to-rule, judge and venue behaviour, on complete coverage of all 94 federal district courts, over ten million cases and 45 million customer-facing documents, by its own figures as of April 2025.

Best for. A federal practice that wants outcome analytics. The blind spot is state depth: the 100 enhanced state courts milestone (May 2024) sits against 200-plus state reporting units.

What customers say. The review footprint is minimal: Capterra carries 2.0/5 (n=1), printed only with its n. The product sells inside LexisNexis enterprise packages, where review platforms rarely see it.

7. Thomson Reuters / Westlaw (incl. CoCounsel, Casetext)

Attribute Detail
Vendor classClass 3 + 2: research platform with derived analytics
Coverage (as published, retrieved 19 Aug 2026)Research corpus plus litigation analytics; Casetext and CoCounsel now sit inside Thomson Reuters
DeliveryPlatform; API access is partnership-gated, with months to onboard absent an existing relationship
Refresh cadenceNot published
Pricing signalPublishes no list price for the configurations a data team would buy
Licensing & redistributionRelationship-gated; not a bulk data supply product
Best forFirms that need research plus derived analytics in one relationship
Not forA data team that needs a feed; this is a research relationship, not a supply contract
What to ask for in the sampleA written statement of what data, if any, may leave the platform
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)Capterra (Westlaw Edge) 4.7/5 (n=21); TrustRadius 8.5/10 (n=39); G2 4.4/5 (n=190)
Evidence qualityThree platforms with real samples; the strongest review base in this list by volume
What reviewers praiseKeyCite for negative-treatment checks; coverage breadth with strong search; a logically organised interface
What reviewers complain aboutProhibitively expensive for solo and small firms; long-term contracts with built-in price increases; charges for documents outside subscription scope
SourceCapterra, TrustRadius, G2 (attributed to the platforms, not the vendor)

What it is. Thomson Reuters / Westlaw is one of the two research incumbents, with Casetext and CoCounsel consolidated inside it. It sells a data buyer a relationship: research, analytics and AI assistance in one contract, with API access gated behind partnership, not a price list.

Best for. An organisation that wants research and analytics from one vendor. Expect months to onboard API access without an existing relationship; do not mistake the platform for a feed.

What customers say. The deepest review base here: Capterra 4.7/5 (n=21), TrustRadius 8.5/10 (n=39), G2 4.4/5 (n=190). Praise: KeyCite and coverage. Complaints: cost mechanics and contract escalators.

8. Bloomberg Law

Attribute Detail
Vendor classClass 3 + 2: research platform with a docket API
Coverage (as published, retrieved 19 Aug 2026)Enterprise legal research corpus with dockets attached
DeliveryPlatform plus a docket API
Refresh cadenceNot published
Pricing signalPublishes no list price; third-party aggregators report figures in the low hundreds per user per month
Licensing & redistributionNot stated publicly
Best forEnterprise research with a docket API attached, if you can absorb an opaque price
Not forTeams that need knowable pricing before a sales conversation
What to ask for in the sampleDocket API terms, volumes and pricing, in writing
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)G2 4.1/5 (n=11); Gartner Peer Insights 4.4/5 (n=6); TrustRadius 8.0/10 (n=3)
Evidence qualityThree platforms, all small samples; directional at best
What reviewers praiseUp-to-date content; complete case-law database with strong docket search; cases, statutes, regulations and legal news in one system
What reviewers complain aboutThe interface could be friendlier and more intuitive
SourceG2, Gartner Peer Insights, TrustRadius (attributed to the platforms, not the vendor)

What it is. Bloomberg Law is the third research incumbent, distinguished by the docket API attached to the research product, and the cleanest demonstration of the category’s pricing opacity.

Best for. An enterprise that wants research plus dockets under one roof. The finding worth printing: Bloomberg Law publishes no list price. Third-party aggregators report low hundreds per user per month; the absent price page is the citable fact, not the estimate.

What customers say. Small samples across three platforms: G2 4.1/5 (n=11), Gartner Peer Insights 4.4/5 (n=6), TrustRadius 8.0/10 (n=3). Praise: content completeness and docket search. The gripe: interface friction.

Category D. Class 2: firm, attorney and legal-entity intelligence

The who layer: neither entry touches a docket; attorney movement, firm economics and entity structure are records no filing feed emits.

9. ALM Intelligence / Law.com Compass

Attribute Detail
Vendor classClass 2: firm and attorney intelligence
Coverage (as published, retrieved 19 Aug 2026)Built on 30+ years of proprietary research: law firm financials, staffing, attorney profiles, client representation, rankings; trusted by more than 100 of the world’s top firms (vendor claims)
DeliveryResearch platform; no API or programmatic delivery advertised on the product website
Refresh cadence“Continuously refreshed through ongoing research cycles” (vendor wording)
Pricing signalPublishes no list price
Licensing & redistributionNot stated publicly
Best forLaw-firm financials, client representation and market structure
Not forTeams that need a programmatic feed; nothing on the product website advertises one
What to ask for in the sampleDelivery format and export rights in writing, since no API is advertised
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)No rating with a published sample size located on G2, Capterra, Gartner Peer Insights, Trustpilot or TrustRadius
Evidence qualityNo published review base on the house platforms; praise exists only on the vendor’s own website, which is marketing, not review data
What reviewers praiseNot printable from review platforms
What reviewers complain aboutNot printable from review platforms
SourceSearched G2, Capterra, Gartner PI, Trustpilot, TrustRadius (19 Aug 2026)

What it is. ALM Intelligence‘s Law.com Compass is the canonical answer to “who are the firms and who represents whom”: three decades of proprietary research into firm financials, staffing and client representation. It appears on zero competing ranking pages.

Best for. Competitive intelligence on law-firm economics and market structure. The structural limitation is delivery: no API is advertised; this is a research product a human reads, not a feed a pipeline consumes.

What customers say. No published review base exists on the five house platforms (19 August 2026). The praise on ALM’s own website is marketing, not review data.

10. Leopard Solutions (SurePoint Legal Insights)

Attribute Detail
Vendor classClass 2: firm and attorney intelligence
Coverage (as published, retrieved 19 Aug 2026)Law Firm Index and Alumni Tracker covering attorney growth and decline, diversity, firm financials, revenue and profits per partner, promotions and retention (vendor claims)
DeliveryBI platform
Refresh cadenceUpdated twice a week (vendor’s stated cadence)
Pricing signalPublishes no list price
Licensing & redistributionNot stated publicly
Best forAttorney movement, lateral tracking and firm headcount
Not forAnything docket-shaped; there is no filing layer here
What to ask for in the sampleContract entity naming: the rebrand is mid-flight, so confirm which entity your paperwork names
Attribute Detail
Ratings (with n, accessed 19 Aug 2026)No rating with a published sample size located on G2, Capterra, Gartner Peer Insights, Trustpilot or TrustRadius (a Glassdoor 3.9/5, n=27, is employee reviews, not product reviews, and is not printed)
Evidence qualityNo published product review base on the house platforms
What reviewers praiseNot printable from review platforms
What reviewers complain aboutNot printable from review platforms
SourceSearched G2, Capterra, Gartner PI, Trustpilot, TrustRadius (19 Aug 2026)

What it is. Leopard Solutions, now a SurePoint company, tracks the who layer’s moving parts: its Law Firm Index and Alumni Tracker are updated twice a week, a rare specific refresh commitment in a category that says “regularly”; hold other vendors to that benchmark.

Best for. Lateral tracking, attorney movement and headcount analysis. The vendor’s own website warns that leopardsolutions.com will soon redirect to SurePoint.com (verified 19 August 2026); two ranking directories still carry the old entity. Confirm which entity your paperwork names.

What customers say. No published product review base exists on the house platforms. The only rating in circulation is a Glassdoor employee score: a workplace measure, not a dataset measure.

The master comparison table

Every legal data provider above, on the seven axes that decide a shortlist (vendor claims, retrieved 19 August 2026).

Provider Class Coverage (as published) Delivery Pricing model Best for Blind spot
Forage AI 4 Scoped per engagement; 500M+ websites crawled Managed feed Scoped, no list price Uncatalogued jurisdictions No catalogue to browse
UniCourt 1 Not verifiable; vendor pages bot-gated API No list price; contract limits Building on a normalised feed Limits unpublished
Trellis Research 1 45 states (vendor support KB) Web platform No list price State trial courts Documents county-by-county
Docket Alarm 1 + 3 Federal + state dockets Platform + alerts $99/user/mo; PAYG tier Tracking at a known price State gaps; PACER on top
Free Law Project 1 (open) 9M+ decisions, 2,000+ courts API + bulk Free; retrieval billed Federal prototyping No derivatives; federal-only
Lex Machina 3 94/94 districts (April 2025) Platform; gated API No list price Federal outcome analytics ~100 state courts
TR / Westlaw 3 + 2 Research + analytics corpus Platform; gated API No list price One-relationship research Not a feed; slow API access
Bloomberg Law 3 + 2 Research corpus + dockets Platform + docket API No list price Enterprise research + dockets Pricing opacity
ALM / Compass 2 30+ years proprietary research Platform, no API No list price Firm financials, market structure No programmatic delivery
Leopard / SurePoint 2 Firm index, twice-weekly refresh BI platform No list price Attorney movement No docket layer
All coverage and pricing figures are vendor claims as published, retrieved 19 August 2026.

Quick Summary

Q: Who are the biggest legal data providers, and which one should I shortlist?

A: The ten here sort into four classes rather than one leaderboard. For federal and state dockets: UniCourt, Trellis, Docket Alarm and the Free Law Project’s open corpus. For outcome analytics: Lex Machina, Thomson Reuters and Bloomberg Law. For firm and attorney intelligence: ALM’s Law.com Compass and Leopard Solutions, now SurePoint Legal Insights. For a jurisdiction no catalogue covers: custom extraction. Shortlist by class before you compare coverage numbers.

What should you ask for in a vendor sample, and what should you measure?

Nine of seventeen ranking pages say “evaluate coverage and accuracy”; none say how. Here is the One-Week Coverage Test instead: one specified sample request and five measurements, each with a benchmark.

Five-step vertical process for the one-week legal data coverage test: measure completeness against court annual reports, filing-to-availability lag, entity resolution, field completion rate, and contractual limits in writing.

The five-field sample request

Request a sample containing case number, court identifier, filing date, parties with roles, and disposition, for a jurisdiction and date range you choose. If the vendor picks the jurisdiction, you measured their best case and learned nothing.

Measuring completeness: 20 known matters, then the ±10% diff

Pick 20 matters you know exist and count the hits. Then copy the Legal Services Corporation‘s Civil Court Data Initiative: diff county-level filing counts against the courts’ own annual reports, a benchmark from counties in over 30 states that flags deviations of more than 10% from the multi-year average. A structured framework for testing a vendor data sample makes this repeatable.

Measuring freshness: filing date to availability date

“How often is the data updated” is unmeasurable; filing-to-availability lag is: the percentage of new filings available on filing day. The benchmark comes from Courthouse News Service v. Schaefer (4th Cir., decided 24 June 2021): measured in May 2018, Norfolk City Circuit Court managed 19% same-day and Prince William County 42.4%; after the litigation, 92.3% and 88.1%-plus. “We update daily” describes the vendor’s job schedule, not the clerk’s.

19% to 92.3% same-day availability in Norfolk City Circuit Court, before and after federal litigation; Prince William County moved from 42.4% to 88.1%-plus. Source: Courthouse News Service v. Schaefer, 4th Cir. 2021 (measured May 2018).

Measuring entity resolution: does the firm resolve to a parent?

Check three joins: firm to parent entity, attorney across firm moves, party to a corporate-registry identifier; the same logic drives resolving firms to parent entities in firmographic work.

Measuring the parse, not the search

An API returning 200 matching dockets says nothing about whether parties, judge and disposition parsed correctly. Search recall is not field quality: measure field completion rate per run, not hit count. And get your contractual limits in writing before signature; in this category they are contract terms, not published figures. Here is what a court-data API returns at a threshold that appears nowhere public:

{ "object": "Exception", "code": "UN429", "message": "TOO_MANY_REQUESTS",
  "details": "Too Many Requests." }

{ "object": "Exception", "code": "UN203", "message": "LIMIT_REACHED",
  "details": "You have hit API limit for the current billing cycle, please contact support@unicourt.com." }

{ "object": "Exception", "code": "UN203", "message": "LIMIT_REACHED",
  "details": "You already have 20 Cases for tracking. Please contact support to increase this limit." }
What you measure What you ask the vendor for What good looks like What a bad answer sounds like
Completeness 20 known matters in a jurisdiction you pick Hits match, county counts within ±10% of court annual reports “Our coverage is unmatched”
Freshness Filing-to-availability lag on new filings Measured same-day percentage, stated per court “We update in real time”
Entity resolution Firm → parent, attorney across firms, party → registry ID Documented resolution logic with identifiers “Our data is fully enriched”
Parse vs recall Field completion rate per field, per run Per-field completion stats on your sample A hit count offered as quality
Contractual limits Monthly ceiling, concurrency, tracking and document caps, in writing Numbers in the contract before signature “Limits are flexible, talk to support”

Quick Summary

Q: What should I ask a legal data vendor for, and what should I measure on it?

A: Ask for a sample containing case number, court identifier, filing date, parties with roles and disposition, for a jurisdiction and date range you choose, not one the vendor chooses. Then run five measurements: completeness against the court’s own published annual report, filing-to-availability lag as a same-day percentage, entity resolution to a parent entity, field completion rate rather than hit count, and your actual contractual limits in writing. Two courts under federal litigation over access delays topped out near 90% same-day availability, so treat any better claim as something to test rather than believe.

Forage AI promotional banner inviting readers to run the one-week coverage test on a scoped pilot: pick a jurisdiction and receive a five-field sample within one to two weeks.

What each vendor class is structurally guaranteed to miss

Some gaps survive any sample test. Every class has a permanent hole, not a roadmap item, each sourced to the judiciary, a federal rule, or a licence.

Four permanent blind spots by legal data vendor class: docket aggregators miss sealed and paper-only records, directories miss what licensing bodies did not report, analytics inherit source-layer gaps and small samples, custom extraction trades away catalogue speed.

Class 1 misses sealed, redacted and paper-only records, and its coverage numbers describe an index, not document access.

“45 states” of coverage can mean retrievable documents in about three dozen counties: one vendor’s own knowledge base pairs a 45-state count with requestable documents in exactly two counties, Los Angeles and Cook.

Under Federal Rule of Civil Procedure 5.2(c), in Social Security and immigration matters a non-party gets only the docket and the disposition remotely, so no vendor can sell those case files. Judicial Conference policy lists nine categories of federal criminal record that never enter the public file.

Nine categories of federal criminal record never enter the public case file, from unexecuted warrants to presentence investigation reports. Source: Judicial Conference of the United States, Privacy Policy for Electronic Case Files.

The largest open corpus cannot even detect sealed items: absence in the data is not evidence of absence in the court. One corrected point: the territories’ federal district courts run CM/ECF and are on PACER; the genuine gap is their local and commonwealth courts (Puerto Rico, Guam, the US Virgin Islands, the Northern Mariana Islands), outside PACER and the state statistical framework. Administrative and ALJ hearings are patchier still.

Class 2 misses whatever the licensing bodies did not report: the attorney census carries forward last year’s rows when a state does not respond. Class 3 misses everything its source layer missed, plus any venue below its per-judge sample threshold; once you hold the documents, the work becomes parsing them, where automating contract data extraction picks up. Class 4 misses the speed of a catalogue, because scoping precedes data. When a class’s guaranteed gap sits on your question, custom extraction is the answer, and that is where Forage AI operates: it delivers the data, not just the pipeline, on an operating base of 500M+ websites.

This article is for informational purposes only and does not constitute legal advice. Consult a qualified attorney for legal guidance specific to your matter.

Quick Summary

Q: What is each kind of legal data provider structurally guaranteed to miss?

A: Docket aggregators miss sealed, redacted and paper-only records, and their state counts describe an index rather than document access. Attorney and firm directories inherit a national census that carries forward last year’s rows when a state does not respond. Analytics products inherit every gap in their source layer. Under Federal Rule of Civil Procedure 5.2(c), in Social Security and immigration matters no vendor can sell you the case file, because non-party remote access is limited to the docket and the disposition.

Forage AI promotional banner for intelligent document processing: turning court filings, contracts and case files into structured fields with a three-layer QA team.

Pricing, licensing, and what you are actually buying

Coverage decides the shortlist; the contract decides what you pay. Twelve of seventeen competing pages say “custom pricing” and stop. The honest version: one public anchor, Docket Alarm’s $99 per user per month with PACER at cost (as of August 2026), and one cost driver, PACER’s $0.10 per page, underneath every Class-1 product.

$99 per user per month, the one published price anchor in the category, with PACER fees billed separately at cost. Source: Docket Alarm pricing page, as of August 2026.

Pricing model What you actually pay for Where the cost hides Who it suits
Per-page pass-through The court’s own access fees Uncapped name searches Anyone touching PACER directly
Per-seat subscription Platform access per user Contract escalators, out-of-scope charges Research-led teams
Pay-as-you-go per document Each retrieval Repeat charges, per-document accumulation Low-volume docket work
Bulk or enterprise licence Volume access under contract terms Unpublished limits, redistribution clauses Teams building on a feed
Scoped managed feed A maintained pipeline to your schema Scoping time before first data Uncatalogued jurisdictions

Licensing outranks price the moment you build on the feed.

The largest free legal corpus forbids derivatives: the Free Law Project publishes over nine million decisions under Creative Commons BY-ND 4.0, the clause that kills a product built on top of it.

Get the redistribution and derivative-works clause in writing before you scope anything; the same questions apply to any data-as-a-service delivery model. And where no catalogue covers your jurisdiction, the build-versus-buy line moves: maintenance, not the build, is the recurring cost. Forage AI handles selector drift, anti-bot evolution, and schema changes as part of the service.

Quick Summary

Q: How much does legal data cost, and what are you actually buying?

A: Five pricing shapes dominate: per-page pass-through, per-seat subscription, pay-as-you-go per document, bulk or enterprise licence, and scoped managed feed. The one genuinely published anchor in the category is $99 per user per month, with PACER fees billed separately at cost on top. The bigger question is licensing: the largest open legal corpus is published under a no-derivatives licence, so a team building a product on it has a legal problem before it has a technical one.

Which provider fits which job?

“Which legal database is best” appears on five of nine captured results pages, and it is the wrong question; the right one is which class your job belongs to. Read down the job column, then run the sample test on your shortlist.

The job you are trying to do Vendor class Shortlist What to test first
Monitor new filings against a client list 1 Docket Alarm, UniCourt; Courthouse News for same-day alerts Filing-to-availability lag in your courts
Build outcome analytics for a federal practice 3 Lex Machina, Bloomberg Law Per-judge n in your venues
Track attorney movement and firm headcount 2 Leopard/SurePoint, ALM Compass Refresh cadence in writing
Resolve law firms to parent entities and clients 2 OpenCorporates plus a firm-intelligence vendor The firm → parent → registry join
Cover a state or county nobody catalogues 4 Custom extraction (Class 4) A scoped pilot on one jurisdiction
Research and prototype on federal case law 1 (open) Free Law Project Whether the ND licence permits your use
Track bankruptcy filings daily 1 (specialist) Epiq AACER Coverage against your district list
Merge two vendors’ feeds into one schema any two SALI LMSS as the join key Field-level mapping on a shared sample

The specialists earn one line each. Epiq AACER publishes coverage of “93 U.S. bankruptcy courts,” updated daily, back to 2007; the judiciary operates 90, a difference we print rather than resolve. OpenCorporates publishes 140+ jurisdictions from government registries, API and bulk; Premonition claims 325M+ cases across 13 countries (both vendor claims, retrieved 19 August 2026). Docket Navigator owns patent litigation, Courthouse News same-day new-filing alerts, InformData court-record retrieval for background screening. judyrecords publishes 770 million-plus cases and 1,426.0 million structured parties (changelog, 7 August 2026) with no jurisdiction map. A record count with no jurisdiction map is not coverage.

770 million+ cases and 1,426.0 million structured parties, with no published jurisdiction map. Source: judyrecords changelog, 7 August 2026.

Quick Summary

Q: Which legal database is best?

A: There is no single best one, and the question hides the real decision. Match the job to a class first: new-filing monitoring and docket tracking are Class 1, attorney movement and firm financials are Class 2, outcome and motion analytics are Class 3, and an uncatalogued jurisdiction is Class 4. Then shortlist two vendors inside that class and run the sample test on both before you compare their coverage claims.

Expert Insights

The two most useful outside voices here run the two institutions this article leans on: the national state-court statistics archive and the largest open legal corpus.

Nicole Waters, Director of Data, Analytics, and Forecasting, National Center for State Courts (State Justice Institute release, 10 November 2025):

“We’ve been tracking traffic trends for some time, so we’re not surprised by the 2024 data showing that traffic filings continued to recover after one of the largest drops during the pandemic. By tracking these state and national trends, we provide important data-driven insights courts can use in their decision-making.”

Michael Lissner, Co-founder and Executive Director, Free Law Project (Above the Law, 15 October 2024):

“A big part of this is through CourtListener, our free platform that provides access to millions of legal opinions, oral arguments, and court documents. We’ve also worked extensively to open up PACER data, which is usually locked behind a paywall, by archiving and advocating for reforms there. Our tools like the RECAP browser extension make it easy for people to download and share PACER documents.”

Frequently asked questions

Is there a free alternative to PACER?

Close: the RECAP archive and CourtListener hold nearly every federal case. But retrieval bills your own PACER credentials, the Fetch API caps at 30 requests a minute, federal-only, and the licence is CC BY-ND 4.0. “Free” is not “usable in your product.”

What databases do law firms use?

Four different product classes get called legal data: docket feeds, firm and attorney directories, outcome analytics, and custom extraction for what those three do not cover. Research subscriptions like Westlaw, LexisNexis and Bloomberg Law sit across the second and third. Which a firm uses depends on the question it is asking.

Which legal database is best?

There is no single best one, because the four vendor classes answer different questions; match the job to a class, shortlist inside it, and test a sample. On “is there a ChatGPT for legal”: legal AI assistants sit on top of these feeds, not in place of them, and inherit their coverage gaps.

Is court data always public?

No. Sealed, juvenile and expunged matters are excluded, nine categories of federal criminal record never enter the public file, and under FRCP 5.2(c) a non-party in Social Security and immigration matters gets the full record only at the courthouse. Paper-only records are public yet unavailable to any pipeline.

Run the test again next year

The coverage test is not a procurement gate you clear once. Document tiers move, court portals redesign, and vendors rebrand mid-flight; one entry here changed its trading name while we were verifying it. Re-run the five measurements per new jurisdiction and at every renewal, and hold vendors to the measurable question: not “do you cover my state,” but “what percentage of new filings in my counties were available same-day last quarter, and what is my monthly ceiling in writing.”

You now have the three sentences the intro promised: your class, that class’s guaranteed gap, and your sample request. If the numbers surprise you, in either direction, we would genuinely like to hear it, along with the class mismatches you catch. That is how a list like this stays honest.

Last updated: August 2026. Vendor facts, ratings and pricing were retrieved 19 August 2026 and will be re-verified on the next scheduled review of this article.

2026 Edition · Strategic Guide
How to Get Started With Your Data Acquisition Strategy For AI
A strategic guide for data leaders who don’t know where to start.
Most guides about data infrastructure jump to the technical fix. This one starts a step earlier, at the strategy decision. It helps you see where you stand on the data acquisition maturity curve, what your options are, and what to ask before you pick a partner.
5 Data Acquisition Stages
3 Data Solutions
15 Min Read
Download the e-book
Free. Sent straight to your inbox.
We’ll email you the guide. No spam, unsubscribe anytime.
S
Written by
Sai Subramaniam
Data Infrastructure Enthusiast, Forage AI

Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.

Reviewed by the team of experts at Forage AI for accuracy and clarity.

Related Blogs

post-image

AI & NLP for Data Extraction

August 21, 2026

LLM Data Extraction: Why Relevance-Based Extraction Fails in Production

Author name

5 min read

post-image

Uncategorized

August 21, 2026

US Healthcare Provider Data: Every Public Source, Ranked (2026)

Author name

5 min read

post-image

AI Infrastructure and Data Management

August 21, 2026

Parquet vs CSV vs JSON: Choosing the Right Delivery Format

Author name

5 min read

post-image

Data Extraction

August 21, 2026

Rossum Alternatives: 15 IDP Platforms Compared for Document Processing (2026)

Author name

5 min read