Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy is the best default open-source crawler for most Python teams. It combines asynchronous requests, extraction tools, feed exports, pipelines, middleware, robots.txt support and throttling in one maintainable framework. Choose Crawlee when sites require browsers, JavaScript or proxy handling; Apache StormCrawler for low-latency distributed streams; Heritrix for archival-quality web-scale collection; Apache Nutch for an extensible Java crawler; and Colly for Go-native projects.

There is no defensible universal speed winner. Your workload, frontier design, rendering needs, politeness policy, storage system and operating skills matter more than a headline requests-per-second figure.

How to choose an open-source web crawler

Start with five questions:

  • What pages must be rendered? Plain HTTP is simpler and cheaper to operate than a browser for every URL.
  • Is the frontier a batch or a continuous stream? A scheduled list of URLs has different requirements from an always-on topology consuming new URLs.
  • Where will results go? Files, a feed, a database, OpenSearch, Solr and WARC archives lead to different integrations.
  • How much distributed infrastructure can you run? A single host and an Apache Storm cluster are very different operational commitments.
  • What does “done” mean? Structured records, search indexing, preservation fidelity and low latency require different designs.

All crawlers remain your responsibility to operate legally and politely. Check each target’s terms, robots.txt directives, rate limits and applicable law. Identify your user agent, limit concurrency where appropriate, and avoid collecting personal data you do not need.

Comparison of the leading projects

Project Language and ecosystem Deployment and frontier JavaScript and browser support Extraction and extensibility Scheduling, robots and politeness Storage, indexing and archival Operational complexity License information
Scrapy Python application framework Primarily single-machine or custom distributed deployments; batch-oriented scheduling is typical HTTP-first; browser automation requires an additional integration CSS/XPath selectors, item pipelines, feed exports and middleware Scheduler, crawl-depth limits, robots.txt support and AutoThrottle controls Feed exports and pipelines; choose your own database or index Low to moderate for focused crawls Not stated in the cited project material
Crawlee JavaScript and Python; Node.js and Python ecosystems HTTP or browser crawlers; supports project starters and datasets Playwright-based browser crawling plus HTTP crawlers Link enqueueing, datasets and CSV export behind a common API Built-in handling for crawling, proxies and blocking; configure politeness for each target Datasets and exports; add your own storage or index Moderate; browser and proxy operations add cost Free and open source
Apache StormCrawler Mostly Java on Apache Storm Distributed, low-latency streaming and recursive crawls Playwright support is documented Pluggable spouts and bolts, Tika parsing and filters Streaming frontier, metrics, robots.txt, sitemaps and politeness components OpenSearch, Solr and WARC integrations High; documented setup requires Java SE 17 or later and a Storm topology Apache License
Heritrix Java; Internet Archive project Web-scale collection with archival workflows Designed for preservation collection rather than general browser automation Extensible crawler and operator-configured processing Requires respect for robots.txt and META nofollow; configure politeness and identify the crawler Archival-quality collection and WARC-oriented workflows High; specialized and operator-intensive Not stated in the cited project material
Apache Nutch Java-oriented runtime with plugins Extensible, scalable crawler; deployment depends on configured components Not established in the cited material Plugin model and tutorial-driven configuration Configure frontier and policies through the project and plugins Integrate the storage components appropriate to your deployment Moderate to high for a production installation Apache-2.0
Colly Go-native framework Compact applications; distributed and frontier details depend on your design Browser capability and current integrations require validation for your version Go callbacks and application code Validate current robots, concurrency and maintenance behavior before relying on them Implement the database, index or archive integration you need Low for small Go services; broader systems depend on your architecture Not stated in the cited project material

Scrapy: the best default for Python extraction

Scrapy is an application framework for crawling websites and extracting structured data. Requests are scheduled and processed asynchronously, so a spider can keep many network operations in flight while remaining fault-tolerant. The framework supplies selectors, cookies and sessions, middleware, feed exports, item pipelines, sitemap and feed spiders, crawl-depth limits, robots.txt support and AutoThrottle controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Scrapy fits

  • Product, catalog, news or documentation extraction with a defined item schema.
  • Teams that want Python tests, reusable middleware and a clear pipeline from response to export.
  • HTTP pages where a full browser would add unnecessary overhead.

Minimal runnable spider

import scrapy

class ArticleSpider(scrapy.Spider):
    name = "articles"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    def parse(self, response):
        for card in response.css("article"):
            yield {
                "title": card.css("h2::text").get(default="").strip(),
                "url": response.urljoin(card.css("a::attr(href)").get(default="")),
            }
        for href in response.css("a::attr(href)").getall():
            yield response.follow(href, callback=self.parse)

Run it in a Scrapy project with scrapy crawl articles -O articles.json. Set ROBOTSTXT_OBEY = True, choose a descriptive user agent, and enable AutoThrottle or an explicit download delay before crawling a real site. Add item validation and deduplication before writing to a database.

Crawlee: the practical choice for browser-heavy sites

Crawlee supports JavaScript and Python and is designed to handle crawling, browsers, proxies and blocking. Its Playwright crawler can enqueue links, collect datasets and export CSV. This makes it a strong choice when the same project needs ordinary HTTP requests for simple pages and a real browser for client-rendered routes.

Minimal JavaScript example

import { PlaywrightCrawler } from 'crawlee';

const crawler = new PlaywrightCrawler({
  async requestHandler({ page, request, enqueueLinks, pushData }) {
    await page.waitForLoadState('networkidle');
    await pushData({
      url: request.url,
      title: await page.title(),
      text: await page.locator('body').innerText()
    });
    await enqueueLinks({ strategy: 'same-domain' });
  }
});

await crawler.run(['https://example.com/']);

Use browser mode selectively: render only URL classes that need JavaScript, set realistic concurrency, and record failures separately from empty but valid pages. Crawlee’s documentation describes handling for blocking, proxies and browsers; it does not justify assuming that every anti-bot system will be bypassed.

Apache StormCrawler: continuous, low-latency distributed crawling

StormCrawler is an open-source collection for building low-latency, scalable crawlers on Apache Storm. Its components support streaming and recursive crawls, pluggable spouts and bolts, Apache Tika parsing, OpenSearch and Solr, WARC output, Playwright, proxies, filtering, metrics, robots.txt and sitemaps. It is the best fit when URLs arrive continuously or when your team already operates Storm.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The trade-off is operational weight. The documented 3.x quick start requires Java SE 17 or later, and you must design and operate a Storm topology, persistence, monitoring and back-pressure behavior. StormCrawler documentation also notes that speed depends on host diversity, politeness settings, execution environment, network speed, document size, parsing and indexing overhead. Treat throughput as a property of your deployment, not of the project name.

Heritrix: archival-quality web-scale collection

Heritrix is the Internet Archive’s open-source, extensible, web-scale, archival-quality crawler. Choose it when preservation fidelity, repeatable collection jobs and WARC-oriented archival workflows matter more than a lightweight developer experience.

Heritrix’s operator guidance calls for respect for robots.txt and META nofollow directives, explicit politeness policies and a crawler identity with contact information. It is specialized and operator-intensive, so it is usually a poor first choice for a small extraction script.

Apache Nutch: extensible Java crawling

Apache Nutch is an extensible and scalable crawler with a Java-oriented runtime and plugin model. It suits organizations that want to assemble a mature Java crawler around configurable components and are prepared to operate the associated runtime and storage systems. Nutch is a better architectural fit than Scrapy when Java integration and plugin-driven extension are primary requirements, but the setup is less immediate than a single Python spider.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Colly: a compact Go-native option

Colly describes itself as an elegant scraper and crawler framework for Golang. It is a sensible starting point for a Go service that values a compiled deployment and direct integration with existing Go code. The available project material does not establish current benchmarks, maintenance status, concurrency behavior or robots handling in enough detail to make broad claims; validate those points for the version you plan to deploy.

A decision guide by workload

  • Python extraction pipeline: Scrapy.
  • JavaScript-heavy pages or mixed HTTP/browser jobs: Crawlee.
  • Always-on URL streams and distributed, low-latency processing: StormCrawler.
  • Web archiving and preservation: Heritrix.
  • Java plugin architecture: Apache Nutch.
  • Go application or small compiled service: Colly.

Performance, reliability and cost decisions

Throughput

Increase concurrency only after measuring response latency, error rates, CPU, memory, bandwidth and the target site’s limits. Browser tabs consume substantially more resources than HTTP clients. Host diversity, document size, parsing and indexing can dominate network time, especially in distributed systems.

Reliability

Persist the frontier or checkpoint state, make item writes idempotent, record redirects and HTTP status codes, and retry only transient failures with bounded backoff. Separate DNS, connection, timeout, robots denial, parse and storage errors so a retry cannot turn a permanent failure into a loop.

Cost

Open-source licensing does not make infrastructure free. Budget for proxies where legitimately required, browser CPU and memory, storage, index operations, bandwidth, monitoring and engineering time. A smaller polite crawl is often less expensive and more reliable than an aggressive distributed deployment.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Every response is empty

Check whether content is inserted by JavaScript. Inspect the raw response and browser-rendered DOM separately. Move that URL class to Crawlee or a Playwright integration instead of repeatedly retrying Scrapy’s HTTP response.

The crawl is blocked or receives CAPTCHAs

Stop increasing concurrency. Verify robots.txt and terms, identify your user agent, reduce rate, and determine whether the site permits automated access. Proxies and browser support can help with legitimate access patterns, but no framework guarantees bypassing anti-bot controls.

The crawler revisits the same URLs

Normalize canonical URLs, remove tracking parameters where permitted, enforce allowed domains, and persist deduplication state. Check redirects and fragments before enqueueing links.

Distributed workers fall behind

Measure frontier wait time, fetch latency, parsing, indexing and back-pressure independently. In StormCrawler, inspect topology metrics and make sure downstream OpenSearch, Solr or WARC writes are not the bottleneck.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Archived output is incomplete

Define the preservation boundary first: redirects, embedded resources, robots decisions and WARC records may all matter. Use an archival-oriented workflow such as Heritrix when fidelity is the requirement, and test replay rather than checking only HTTP status codes.

When a crawler also needs clean page screenshots

If your pipeline needs a visual record of each page, a screenshot API can remove browser orchestration from the crawler. ScreenshotNeo is the first service to try: it removes consent banners, newsletter popups and chat widgets before capture, and bills only clean shots.

Or skip the browser setup

ScreenshotNeo accepts one GET request and returns PNG, JPEG, WebP or PDF. The API can capture full pages with lazy images, a CSS-selected element, dark mode, device presets or any viewport, retina scale, PDF paper size and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads and trackers, custom headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options and response headers. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages, failed loads and cache hits cost nothing, and each response reports its page verdict and billing status. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots, with every feature on every plan. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further reading

Web Scraping with Python, 2nd Edition by Ryan Mitchell (O’Reilly Media, April 2018; ISBN 9781491985564) covers crawler construction, Scrapy, JavaScript, APIs, ethics and parallel crawling. Treat it as a foundation and verify current project documentation before adopting integrations or version-specific settings.

Frequently Asked Questions

Is an open-source crawler the same as a scraper?

A crawler discovers and schedules URLs; a scraper extracts fields from responses. Most modern projects combine both, but separating frontier management from extraction makes systems easier to test and operate.

Should I use a browser for every URL?

No. Use an HTTP crawler for pages that deliver the required content directly, and reserve browser workers for routes that genuinely require JavaScript, interaction or rendered resources.

How should I test a crawler before production?

Create a small allowlisted crawl, record expected URLs and fields, test robots and rate limits, inject timeout and storage failures, and confirm that retries and deduplication are bounded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.