Alternative Data for Hedge Funds: A Practical Guide (2026)

Every fund has alternative data now. A decade ago, knowing that card spending at a retailer was up before the company reported was an edge in itself. Today it is table stakes, and the edge has moved somewhere harder to copy: sourcing data your competitors cannot get, evaluating it without fooling yourself, and integrating it before the signal is crowded out. Owning a dataset is not alpha. What you do with it is.
This guide is written for the people who actually run that process: quant researchers, portfolio managers, and data-sourcing teams deciding which datasets are worth onboarding and how to turn them into a tradable signal. We are not going to re-explain what alternative data is from scratch (if you want the primer, start with our explainer on what alternative data is), and we are not going to hand you another ranked list of vendors (that lives in our guide to the top alternative data vendors). This is the part in the middle that most articles skip: how a fund chooses, tests, and sources alt data for alpha.
We will walk through why funds use it, the datasets that actually move a position and the signal each one gives, the evaluation method that separates a real edge from an overfit backtest, the compliance lines you cannot cross, and the build-versus-buy decision underneath all of it. Buy-side spending on alternative data now runs into the billions and keeps compounding at double digits, so the question is no longer whether to use it, but how to use it well.
Quick Digest
- Nowcasting: Card, geolocation and web-scraped data estimate what a company is doing now, weeks before the official disclosure lands.
- Point-in-time integrity: If you cannot reconstruct what a dataset looked like on each historical date, look-ahead bias will inflate the backtest.
- Alpha decay: The more funds license the same off-the-shelf panel, the faster the edge is arbitraged away.
- Legal footing: Material non-public information is still MNPI however it is packaged, so licensing, personal data and collection method are diligence questions.
- Sourcing: Marketplaces and direct vendors are fastest for standard panels, while the exclusive datasets usually need custom collection.
Why do hedge funds use alternative data?
The core use is nowcasting. Traditional fundamentals tell you what a company did last quarter, reported weeks after the quarter closed. Alternative data tells you what is happening now: card and transaction data can estimate a retailer’s revenue before the earnings release, geolocation can show whether store traffic is rising, and web-scraped pricing can reveal demand and pricing power in near real time. The fund that sees the number forming has an informational edge over the consensus still waiting on the 10-Q.
The second use is conviction. Discretionary managers use alt data to validate or challenge a thesis, who is hiring, whose app is gaining users, whose suppliers are wobbling, while quant funds fold it into systematic signals. Either way, the value is the same: a view that is earlier, more granular, or more honest than the official disclosure. What separates the funds that profit from it from the funds that just spend on it is uniqueness and speed, not access.
Which datasets actually move a position?
Alternative data is not one thing. Each category answers a different question, and the skill is matching a dataset to a signal you can actually trade. The map below is how practitioners think about it, by the question each dataset answers, not by the vendor that sells it.
Alternative datasets and the signals they generate
| Dataset type | What it tracks | Signal for the fund | Example use |
|---|---|---|---|
| Consumer transaction / card | Spending by merchant or brand | Revenue, market share, churn | Nowcast a retailer’s quarter |
| Geolocation / foot traffic | Visits to stores and sites | Demand, property and REIT activity | Track mall or chain traffic trends |
| Web-scraped (pricing, product, reviews, jobs) | Prices, catalogs, hiring | Demand, pricing power, growth proxy | Detect price moves and hiring surges |
| App downloads & usage | Installs, active users | Consumer-tech growth and engagement | Gauge a consumer app’s momentum |
| Satellite & geospatial | Physical activity | Commodity supply, output, logistics | Read oil storage or parking lots |
| Sentiment / NLP / news | Tone, events, mentions | Event detection, sentiment shifts | Flag supply-chain or earnings signals |
| ESG & supply chain | Supplier links, risk events | Risk exposure, network effects | Spot supplier disruption early |
Web-scraped data deserves a closer look because it is the most customizable category. Prices, product catalogs, customer reviews, and job postings are public, high-frequency, and specific to the names you trade, and a job-posting surge or a quiet price cut often shows up well before it reaches a financial statement. The catch is that the most valuable web data is rarely sold neatly off the shelf, which is exactly where custom sourcing comes in later.

How do you test a dataset for real alpha?
This is the part that decides whether a dataset makes money or just makes a slide. A dataset that looks brilliant in a backtest can be worthless live, usually for one of a handful of reasons. Here is the checklist a disciplined fund runs before it onboards anything.
Point-in-time integrity comes first. Was the data actually available at the timestamp it claims, or has it been revised after the fact? If you cannot reconstruct what the dataset looked like on each historical date, look-ahead bias will inflate your backtest and the signal will evaporate in live trading. This is the single most common way an alt-data backtest lies to you.
Expert Insights
David H. Bailey, Jonathan Borwein, Marcos Lopez de Prado and Qiji Jim Zhu, writing in the Notices of the American Mathematical Society in 2014, showed that "high simulated performance is easily achievable after backtesting a relatively small number of alternative strategy configurations". That is the statistical reason point-in-time evidence matters more than a headline Sharpe ratio: if a researcher never says how many configurations were tried, the backtest cannot tell you how much of the result is selection.
Then uniqueness and alpha decay. How many other funds already license this exact dataset, and how fast does the edge erode as it gets crowded? A widely-sold panel is largely priced in; a raw or exclusive source holds its alpha longer. Weigh that against history depth (enough data to backtest across more than one regime), coverage and entity mapping (does it map cleanly to the tickers you trade), panel stability (methodology changes silently break signals), latency (timely enough for your holding period and your pipeline), and legal cleanliness, which we treat separately below because it is non-negotiable.
The alternative-data evaluation checklist
| Criterion | The question to ask | Why it matters |
|---|---|---|
| Point-in-time integrity | Was the data available at the timestamp it claims? | Kills look-ahead bias in backtests |
| History depth | Enough history to test across cycles? | Credibility of the signal |
| Alpha decay / crowding | How many funds have it; how unique is it? | How long the edge survives |
| Coverage & entity mapping | Does it map cleanly to your tickers? | Whether you can actually use it |
| Quality & panel stability | Is the methodology consistent over time? | Signal stability |
| Latency & delivery | Timely for your horizon and pipeline? | Operational fit |
| Legal cleanliness | Licensing, PII, collection, usage rights? | Compliance and risk |
A backtest without point-in-time data is a story, not evidence. If you cannot confirm what the data looked like at each historical timestamp, look-ahead bias will quietly inflate the signal, and it will not survive contact with live capital. Demand point-in-time history before you demand performance.
The compliance lines a fund cannot cross
This article is for informational purposes only and does not constitute legal advice. Consult a qualified attorney for legal guidance specific to your situation.
Alternative data is not exempt from securities law. The packaging does not change the substance: material non-public information is still MNPI whether it arrives as a tip or as a dataset, and the history of expert-network and data-licensing enforcement is the cautionary tale every fund’s compliance team already knows. The rule of thumb is simple to state and harder to operationalize: use non-material, legally-sourced data, license it cleanly, and know how it was collected.
That last point matters most for web data. Extracting public web information has been broadly upheld, but it is governed by site terms and privacy law, so the collection method and the handling of any personal data are part of your diligence, not the vendor’s problem alone. Funds run vendor-risk and data-diligence reviews for exactly this reason: licensing terms, PII exposure, collection methodology, and usage rights all sit on the compliance checklist before a dataset reaches a researcher.
Expert Insights
The SEC Division of Examinations, in its 2022 risk alert on investment adviser MNPI compliance issues, reported that "Advisers did not appear to adequately memorialize diligence processes or follow them consistently and instead engaged in ad hoc and inconsistent diligence of alternative data service providers". Writing the diligence down is part of the control, not paperwork after the fact.
Material non-public information is still MNPI no matter how it is packaged. Alt data must come from non-material, legally-obtained sources with clean licensing and a defensible collection method. Treat collection method and PII as first-class diligence questions, and prefer providers who can document both.

Build vs buy: sourcing alternative data
Once you know what you want and how to judge it, the question is how to get it. There are three routes, and most funds use all three depending on the dataset. The decision usually comes down to one thing: exclusivity.
Marketplaces and aggregators (Nasdaq Data Link, BattleFin, Eagle Alpha, Neudata) are where you discover and trial datasets quickly. Direct vendors sell a specific, productized panel, fast to onboard, but also sold to everyone else. Managed or custom web extraction is how funds get the exclusive datasets that are not on any shelf: a competitor’s full pricing history, a sector’s hiring velocity, a supplier network mapped to your tickers, collected to your specification and delivered point-in-time. The trade is build effort and data-ops cost against a longer-lived, less-crowded edge.
Where funds source alternative data, by type (a starting map, not a ranking)
| Dataset type | Example sources | Sourcing route |
|---|---|---|
| Consumer transaction | Earnest Analytics, Consumer Edge, Facteus | Direct / marketplace |
| Geolocation | Advan, Placer.ai | Direct |
| Web-scraped / custom | Forage AI (managed/custom), YipitData, Thinknum | Managed extraction / vendor |
| Satellite | Orbital Insight, RS Metrics | Direct |
| Sentiment / NLP | RavenPack | Direct |
| Discovery / marketplaces | Nasdaq Data Link, BattleFin, Eagle Alpha, Neudata | Aggregator |
This table is a starting map, not a verdict. For a fuller view of vendors across every category, see our guide to the top alternative data vendors.
Bought-by-everyone data decays fastest. The more funds license the same off-the-shelf panel, the quicker the alpha is arbitraged away. Exclusive or custom-sourced data costs more to stand up, but it is the part of the sourcing budget that keeps working after the crowd arrives.
Where managed web data extraction fits
Forage AI is the managed-acquisition option in that sourcing decision. We are not a marketplace and not an off-the-shelf panel. We build and run custom web data extraction for funds: the exclusive pricing, product, review, and hiring data tied to the entities you trade, collected compliantly, structured to your schema, and delivered point-in-time into your research pipeline. You define the dataset and the universe; we own the extraction, the anti-detection, the entity mapping, and the maintenance, so your researchers spend their time on the signal, not on scraping infrastructure.
Frequently asked questions
How do hedge funds use alternative data?
Mainly to nowcast company metrics ahead of official disclosure, transaction data to estimate revenue, geolocation to read demand, web data to track pricing and hiring, and to validate or challenge an investment thesis with granular, real-time evidence. Quant funds fold it into systematic signals; discretionary funds use it for conviction. The goal is a view that is earlier or more accurate than consensus.
What is point-in-time data and why does it matter?
Point-in-time data preserves what a dataset actually looked like on each historical date, before any later revisions. It matters because without it a backtest can use information that was not available at the time, called look-ahead bias, which inflates the apparent signal. A dataset you cannot reconstruct point-in-time should be treated as unproven, however good its backtest looks.
What is alpha decay in alternative data?
Alpha decay is the erosion of a signal’s edge as more funds trade on the same data. A widely-licensed dataset gets arbitraged toward efficiency, so its alpha fades. Unique or custom-sourced data decays more slowly, which is why exclusivity is a core part of how funds evaluate and source datasets.
Is alternative data legal for hedge funds?
It is legal when the data is non-material, legally obtained, and properly licensed. The risks are material non-public information, mishandled personal data, and improper collection methods. Extracting public web data is broadly permissible but is governed by site terms and privacy law, so funds run vendor diligence on licensing, PII, and collection method before onboarding a dataset.
Should a fund build its own data pipeline or buy datasets?
Both. Marketplaces and direct vendors are fastest for standard panels, but the exclusive datasets that hold their alpha usually require custom collection. Building web extraction in-house carries a real, ongoing data-ops burden, so many funds use a managed extraction partner for the custom datasets and buy the commoditized ones, reserving engineering effort for signal research.
Related reading
- What Is Alternative Data? A Practical Guide
- Evaluating Top Alternative Data Vendors
- Top Financial Data Providers 2026
- Ecommerce Data Scraping: Prices, Reviews & Catalogs
S
Written by
Sai Subramaniam
Data Infrastructure Enthusiast, Forage AI
Sai is a data infrastructure enthusiast who has spent the past two to three years following the AI space closely, from the infrastructure layer to the fast-growing world of data for AI. He is genuinely curious about how modern data pipelines get built and where the data industry is heading, and he writes insightful pieces on the core topics that shape this niche.
Reviewed by the team of experts at Forage AI for accuracy and clarity.