Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hedge funds use web scraping to collect and organize web-derived observations—such as product prices, reviews, app-store information, and public website activity—that may help them assess a company or market. These observations are research inputs, not automatic trading signals: they need to be checked for provenance, representativeness, consistency, legal and privacy risks, and relevance to a specific investment question. Scraped web content is also only one part of alternative data, which can include information gathered through other methods.

What “web scraping for alternative data” means

Web scraping is the automated collection of information from web pages or other web-accessible sources. In an investment-research setting, a team may collect changing observations repeatedly, place them in a structured dataset, and compare those observations over time or across companies. A fund can build collection tools internally or buy data from a provider that has collected, cleaned, aggregated, or modeled information.

The resulting data may be described as alternative data because it is not part of a company’s conventional financial statements or other traditional financial sources. The SEC used that description in its September 29, 2021 release about App Annie. But “alternative data” is a broad category, not a synonym for scraped web pages: SEC-filed adviser materials also list transaction, geolocation, satellite, and other data that may be collected by different means.

It helps to separate three layers that are sometimes blurred together:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Observation: what a source displayed or recorded, such as a listed price or a public review.
  • Dataset or estimate: observations cleaned, combined, aggregated, or modeled by a fund or vendor.
  • Investment interpretation: what a researcher infers about demand, competition, or business activity, and whether that inference matters to a decision.

An observation is not proof of a company’s sales, and a vendor estimate is not the same thing as the underlying records. Each layer adds assumptions and possible error.

What kinds of information might be collected?

SEC-filed adviser materials list examples of alternative data including website usage, mobile-app and app-store analytics, public social posts, online browsing activity, product reviews, price trackers, shipping receipts and trackers, and internet activity or quality data. Other listed categories include geolocation or foot traffic, credit-card transactions, email receipts, point-of-sale data, and satellite imagery. These examples describe possible data categories; they do not establish that every dataset is available, lawful to acquire, representative, or useful for every issuer.

Possible input Question it could help investigate Important limitation
Product reviews Are customers discussing product quality, availability, or service? Reviewers may not represent the customer base; platform rules and collection rights still matter.
Online prices and product listings Are listed prices, promotions, or apparent product availability changing? A listed price or listing status does not establish completed sales or inventory.
Website or app-use measures Is digital engagement changing, according to the source or its model? Coverage, methodology, and the difference between measured usage and an estimate need scrutiny.
Shipping or internet activity information Could external activity provide context on a company or sector? The reviewed materials do not establish a validated hedge-fund trading signal from these inputs.
Geolocation, transactions, or satellite data Could a non-web observation complement a web-derived measure? These are distinct data categories and can involve separate collection, privacy, and compliance risks.

These are analytical possibilities, not documented strategies with established returns. A signal that appears intuitively relevant may still have gaps, delays, selection bias, or a weak relationship to the business question.

How a fund can turn a web observation into research

There is no single workflow established for every hedge fund. A sensible research process, consistent with the controls described in SEC-filed adviser materials, starts with the question rather than with whatever data happens to be available.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define the question. State what business activity the team wants to understand—for example, whether consumer interest in a product may be changing—and what observable evidence could plausibly bear on it.
  2. Choose a source and collection route. Decide whether to collect internally or evaluate a provider. Record what source is involved, who collects it, and whether the output is a raw observation, an aggregation, or a modeled estimate.
  3. Establish permission and controls before collection. Review the access method, applicable terms and contracts, privacy exposure, site impact, and the firm’s approval requirements. Do not treat public visibility as a complete legal or compliance analysis.
  4. Assess the data itself. Check coverage, update frequency, historical depth, missing observations, consistency, latency, and whether changes in the source or vendor methodology can be identified.
  5. Test whether it answers the question. Compare the measure with other relevant evidence and consider alternative explanations. A plausible proxy is not automatically a useful or predictive one.
  6. Monitor the source and use. Keep records of approvals, permitted uses, provider representations, and material changes. Revisit the assessment if collection methods, data fields, or vendor practices change.

The available SEC-filed materials describe categories and controls; they do not demonstrate that this process produces a general return premium or that any particular web signal predicts investment performance.

Build collection internally or buy from a provider?

Internal collection can give a team more direct knowledge of how observations are gathered and transformed. It also makes the firm responsible for the collection process, operational behavior, privacy controls, documentation, and maintenance. A provider may offer cleaned or modeled outputs, but outsourcing collection does not remove the need to understand provenance, rights, methodology, and permitted use.

When evaluating either route, examine the same core questions:

  • Provenance and rights: Which sources feed the dataset? Can the collector explain the chain of collection and document permission or a license where needed?
  • Access method: Is collection limited to public areas, or does it involve accounts, access controls, or other restricted material? What permission supports the method?
  • Privacy and sensitive information: Could the data include personal information or material nonpublic information (MNPI)? What minimization, anonymization, aggregation, and escalation controls apply?
  • Site impact and traceability: How are request volume and service impact managed? Can the collector be identified and its conduct reviewed?
  • Data quality: What are the coverage, update cadence, historical depth, known gaps, and methodology for derived estimates?
  • Contract and ongoing oversight: What uses are permitted? How are changes to sources or methods communicated, and how often is diligence refreshed?

SEC-filed alternative-data policies describe provider diligence, documentation, escalation of suspected MNPI or personal information, and review of provider controls. Another SEC-filed code requires compliance pre-approval for new alternative-data providers and products and calls for review of controls intended to prevent MNPI. Those are examples of firm policies, not universal rules or a complete statement of legal obligations.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What compliance controls do SEC-filed policies describe?

One adviser policy filed with the SEC describes pre-approval for scraping projects. Its examples include limiting collection to public portions of websites, avoiding logins and CAPTCHAs unless permission is granted, not disguising the scraper’s identity, avoiding excessive requests that could affect site operation, and avoiding or promptly anonymizing captured personal information.

The policy reflects that firm’s controls, not an SEC rule or a universal safe harbor. Its treatment of unaffirmed embedded terms is likewise a firm-specific policy position; it should not be recast as a general legal conclusion. A separate SEC-filed alternative-data policy describes diligence on whether collection is lawful and consistent with industry standards, whether it is confined to public areas absent a license, and whether it avoids disrupting sites and keeps the collector traceable.

For a particular collection or use, assess the actual source, access conditions, contracts, data type, and jurisdiction with qualified counsel and the firm’s compliance staff. Public availability may be relevant, but it does not by itself resolve every question involving terms, access controls, privacy, intellectual property, contractual restrictions, or other law.

What the App Annie enforcement case illustrates

On September 29, 2021, the SEC announced charges against App Annie and its founder. According to the SEC’s release, App Annie sold app-performance estimates to subscribers, including trading firms. The SEC found that App Annie used non-aggregated and non-anonymized confidential app-performance data to alter model-generated estimates, contrary to its representations about aggregation and anonymization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The case is a concrete reason to scrutinize a provider’s collection rights, underlying inputs, transformations, and representations—not just the apparent usefulness of a finished estimate. It is not evidence that all alternative-data providers or scraped data are unlawful, and it does not establish that every hedge-fund customer was charged or knowingly involved.

Does the hiQ decision make public web scraping legal?

No universal conclusion follows from the Ninth Circuit’s April 18, 2022 hiQ opinion. It concerned a specific dispute over publicly accessible LinkedIn profile data and the Computer Fraud and Abuse Act. It is jurisdiction- and fact-specific, not blanket permission to scrape any site or use any data for investment research. Other circumstances—including access controls, site terms, privacy laws, intellectual-property claims, and contractual restrictions—may matter. The legal status of a particular technique, source, data category, or use cannot be settled by the fact that a page is visible without logging in.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where ScreenshotNeo fits—and where it does not

ScreenshotNeo is a website screenshot API and MCP server for developers. A screenshot can preserve what a page looked like at capture time, but an image is not a structured dataset of prices, reviews, or activity, and using a screenshot service does not settle whether a collection is permitted. Treat it as a page-capture tool, not as a substitute for source diligence, compliance review, or extracting and validating the observations your research question requires.

For a permitted page capture, this single GET request returns an image or PDF; the API documentation lists the available options and response behavior: ScreenshotNeo API docs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

These examples capture a page; they do not parse the image into fields or make a page’s content representative of a market. Use an API key from your account, and check the response and documentation before treating a returned file as a successful capture.

Or skip the browser setup

ScreenshotNeo can capture a page through one API call. Before the capture it accepts the cookie or consent banner as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, or other MCP clients. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. All features are available on every plan.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the API documentation for parameters and response details. Sign up for 1,000 free screenshots a month, with no card required.

Common problems to check before relying on a dataset

  • The page is accessible, so the collection must be cleared: visibility alone does not settle permission or legal questions. Review the access method, terms, contracts, data type, and applicable jurisdictions with counsel and compliance.
  • A vendor calls an output anonymized or aggregated: request an explanation of the source inputs and transformations, the controls behind those representations, and the contractual rights to provide and use the data.
  • A time series changes abruptly: investigate source coverage, collection changes, outages, methodology revisions, and missingness before interpreting the shift as a business event.
  • A metric seems to confirm the investment thesis: check whether it measures the stated question, whether its sample is representative, and whether other explanations fit the observation.
  • Collection causes site impact or encounters a login, CAPTCHA, or other barrier: do not evade the barrier or increase request pressure as a workaround. Stop and obtain appropriate permission or use a properly licensed source.
  • Personal or potentially nonpublic information appears: follow the firm’s escalation and handling procedures; do not assume that a general dataset label makes sensitive contents safe to use.

What the evidence can—and cannot—support

SEC-filed adviser materials document a range of alternative-data categories and examples of firm controls. The SEC’s App Annie action documents a specific dispute over vendor data practices and representations, while the hiQ opinion addresses a specific legal dispute. These sources do not quantify hedge-fund adoption of web scraping, establish a general alpha or return advantage, or validate a particular scraped signal. Any claim about a dataset’s investment usefulness therefore needs evidence specific to that dataset, question, time period, and method.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Is alternative data always scraped from the web?

No. Alternative data includes categories such as transactions, geolocation, and satellite imagery that may be collected through methods other than web scraping.

Does a provider’s estimate mean the fund is using raw scraped observations?

Not necessarily. A provider may clean, aggregate, or model underlying information; diligence should distinguish the source records from the delivered estimate.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.