Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy Splash is a two-part system: scrapy-splash is the Scrapy integration, while Splash is a separate HTTP service that renders pages in a WebKit browser. Install the Python package, run a Splash server (Docker is the usual route), configure the documented middleware and request fingerprinter, then choose a rendering endpoint. Use render.html or render.json for straightforward pages; use /execute or /run when Lua must control navigation, JavaScript, cookies, or returned data.

This guide covers a working setup, Lua scripts, sessions, POST requests, compatibility limits, troubleshooting, and an alternative when maintaining a browser-rendering service is more work than you want.

How the Scrapy Splash architecture works

A normal Scrapy request downloads the response directly. A Splash request sends the target URL and rendering arguments to the Splash HTTP server. Splash loads the page in its WebKit engine, executes JavaScript, and sends the rendered result back to Scrapy.

  • Scrapy: your spider, scheduling, parsing, retries, pipelines, and item extraction.
  • scrapy-splash: Scrapy request classes, middleware, cookie handling, argument de-duplication, and request fingerprinting.
  • Splash: the separately deployed rendering service and its HTTP API.

Installing scrapy-splash alone does not install or start Splash. Your Scrapy process must be able to reach the server address configured in SPLASH_URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prerequisites and installation

Use a supported Python environment

Current Scrapy installation guidance requires Python 3.10 or newer (CPython or PyPy) and recommends a dedicated virtual environment. Create one before installing project dependencies:

python -m venv .venv
# Linux/macOS
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
python -m pip install --upgrade pip
python -m pip install scrapy scrapy-splash

Run Splash with Docker

The commonly documented container command publishes Splash on port 8050:

docker run -p 8050:8050 scrapinghub/splash

Leave that process running. From the host, the service address is usually http://127.0.0.1:8050. If Scrapy runs in another container, do not use 127.0.0.1 to refer to the Splash container; use the Docker service name or another address reachable from the Scrapy container.

Configure Scrapy settings

Set the Splash URL and add the middleware in the documented order. The compression middleware must have a lower priority number than the Splash middleware so Splash responses are handled correctly:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# settings.py
SPLASH_URL = 'http://127.0.0.1:8050'

DOWNLOADER_MIDDLEWARES = {
    'scrapy_splash.SplashCookiesMiddleware': 723,
    'scrapy_splash.SplashMiddleware': 725,
    'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}

SPIDER_MIDDLEWARES = {
    'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}

REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter'

SplashDeduplicateArgsMiddleware avoids repeatedly transferring identical large arguments. SplashRequestFingerprinter makes Splash arguments part of request identity, preventing Scrapy from treating materially different render requests as duplicates.

Choose the right Splash endpoint

render.html and render.json

Use these endpoints when you need a rendered document with little custom control. They accept normal Splash arguments such as a URL, viewport, wait time, or JavaScript settings. render.html returns HTML; render.json wraps rendered output and metadata in JSON.

/execute and /run

The Splash API describes execute and run as its most versatile endpoints because they execute arbitrary Lua rendering scripts. Choose them when you need to click, wait for a condition, evaluate JavaScript, manage cookies, make several navigations, or return a custom table rather than a standard response.

In scrapy-splash, an execution request normally uses endpoint='execute' and passes the script as lua_source. The /run endpoint is also Lua-based; use the endpoint that matches the response format and control flow required by your project.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A minimal Scrapy spider using Splash

This spider renders a page, waits briefly for client-side content, and parses the resulting HTML:

import scrapy
from scrapy_splash import SplashRequest

class ProductSpider(scrapy.Spider):
    name = 'products'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/products']

    def start_requests(self):
        for url in self.start_urls:
            yield SplashRequest(
                url,
                self.parse,
                endpoint='render.html',
                args={
                    'wait': 2,
                    'timeout': 90,
                },
            )

    def parse(self, response):
        for title in response.css('h2.product-title::text').getall():
            yield {'title': title.strip()}

The timeout argument is a Splash-side limit for rendering. Scrapy’s own download timeout and retry settings still apply to the HTTP request made to Splash, so configure both layers for slow sites.

Writing Lua for /execute

The basic pattern

A Lua script defines main(splash). Navigate with splash:go, wait or evaluate JavaScript as needed, and return a string, value, or table:

script = r'''
function main(splash)
    assert(splash:go(splash.args.url))
    splash:wait(2)
    return splash:evaljs("document.title")
end
'''

yield SplashRequest(
    'https://example.com',
    self.parse_title,
    endpoint='execute',
    args={'lua_source': script},
)

def parse_title(self, response):
    # An execute response containing a scalar is available as response.text.
    self.logger.info('Rendered value: %s', response.text)

splash.args.url receives the URL supplied by the Scrapy request. assert turns a failed navigation into a Lua traceback that you can inspect in Scrapy logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Return HTML and a value together

Returning a Lua table is useful when the spider needs both the rendered document and a computed value:

function main(splash)
    assert(splash:go(splash.args.url))
    splash:wait(1)
    return {
        title = splash:evaljs("document.title"),
        html = splash:html()
    }
end

Parse the JSON response according to the endpoint’s response format, and validate that the keys you expect are present before extracting data.

Waiting and JavaScript

A fixed splash:wait(seconds) is simple but can waste time or finish too early. For a page whose content appears after a known script action, call splash:evaljs and then wait for the resulting network and DOM work. If the site needs a browser feature WebKit does not implement, no amount of waiting will make that feature compatible.

Cookies, sessions, and request state

Splash is stateless per request. A later request does not automatically inherit cookies from an earlier one. To preserve a login or shopping session, pass cookies into Lua, initialize them before navigation, and return the updated cookie jar:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
function main(splash)
    splash:init_cookies(splash.args.cookies)
    assert(splash:go(splash.args.url))
    return {
        cookies = splash:get_cookies(),
        html = splash:html()
    }
end

On the Scrapy side, use the session support provided by scrapy-splash, including a consistent session_id, so related requests use the same logical cookie flow. Treat session identifiers as isolated per user or crawl; sharing one identifier across unrelated jobs can mix authentication state.

POST requests and cached Lua arguments

POST support

Splash 1.8 or newer is required for the http_method and body POST arguments. With /execute, the Lua script must pass those values to splash:go rather than assuming a GET:

function main(splash)
    assert(splash:go{
        url = splash.args.url,
        http_method = splash.args.http_method,
        body = splash.args.body,
    })
    return splash:html()
end

Set the method and body in the Scrapy request arguments, and ensure the target site also receives any required content type or authentication headers.

Cached arguments

Splash 2.1 or newer supports server-side caching of large static arguments such as lua_source. This can reduce request traffic and duplicate disk-queue data when many requests reuse the same script. It does not cache the target page’s changing content; it caches request arguments handled by Splash.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compatibility: what can fail and why

Scrapy and Python

Use Python 3.10 or newer for current Scrapy installations, and verify that the scrapy-splash version you select supports the Scrapy release in your environment. Scrapy’s compatibility policy calls out backward incompatibilities in release notes and generally retains deprecated features for at least one year, but that policy does not guarantee that an unmaintained integration will track every new Scrapy release.

Rank #4
ScrapTherapy® Cut the Scraps!: 7 Steps to Quilting Your Way through Your Stash
  • Country of Origin:US
  • CPSIA:N
  • Hazardous?:No
  • Tariff:4901990050

Splash and WebKit

The major compatibility boundary is the browser engine. Splash uses WebKit, and its FAQ warns that target sites can be incompatible with that engine. Modern sites may require browser APIs, JavaScript syntax, anti-bot behavior, or multi-window interaction that WebKit cannot provide.

Scrapy’s dynamic-content guidance positions Splash for JavaScript-rendered pages, while a modern headless browser may be necessary for on-the-fly DOM interaction or multiple windows. Choose Splash when its lightweight HTTP/Lua model fits the site; choose a current browser automation stack when the site’s behavior depends on newer browser capabilities.

Troubleshooting checklist

Connection refused or timeout before rendering

  • Confirm the container is running with docker ps.
  • Open the Splash address from the same network namespace as Scrapy.
  • Check SPLASH_URL; container-to-container setups usually require a service name, not 127.0.0.1.
  • Raise Scrapy and Splash timeouts only after confirming the target page is reachable.

Duplicate requests or unexpected filtering

Ensure SplashDeduplicateArgsMiddleware and REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter' are present. Without the Splash fingerprinter, requests that differ only in rendering arguments can collide with ordinary Scrapy fingerprints.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compressed or malformed responses

Check middleware priorities against the documented values. SplashMiddleware must run before HttpCompressionMiddleware in the effective order shown in the setup block.

Lua traceback or missing HTML

  • Inspect the complete request, endpoint, and Lua traceback.
  • Verify that splash:go is wrapped in assert so navigation failures are visible.
  • Confirm every value used in Lua is supplied in args.
  • For POSTs, verify Splash is 1.8 or newer and that the Lua call passes the method and body.

JavaScript content never appears

Increase the wait only if the page is genuinely slow. Then test whether the site depends on APIs unsupported by Splash’s WebKit engine. A modern headless browser may be required for complex interaction, newer browser features, or multiple windows.

Login state disappears

Remember that Splash requests are stateless. Initialize incoming cookies with splash:init_cookies, return splash:get_cookies(), and keep a stable scrapy-splash session_id for the sequence that must share state.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational and cost considerations

Self-hosting Splash gives you control over deployment and request volume, but you operate the container, networking, logs, retries, and upgrades. Keep verbose logging available while diagnosing failures: run the container with -v2 and inspect the full request and traceback, then reduce verbosity for routine operation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Scrap Quilt Secrets: 6 Design Techniques for Knockout Results
  • Suitable for all kinds of project works
  • Acid and toxic free
  • Designed for easy usage

Rendering is more expensive than downloading raw HTML because Splash runs a browser engine for each request. Reuse Lua scripts, enable cached arguments where supported, avoid unnecessary waits, and request only the endpoint output your parser needs. There are no authoritative performance benchmarks in the documented material, so size capacity with measurements from your own target sites rather than a published throughput number.

Or skip the browser setup

If you need screenshots or PDFs rather than a self-managed Splash rendering service, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled.

Only clean shots are billed. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A direct call looks like this:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free, and every feature is available on every plan. Sign up free for ScreenshotNeo and start with the 1,000 no-card screenshots.

Frequently Asked Questions

Can I run Splash without Docker?

Yes. Docker is the commonly documented deployment, but Splash is a separate HTTP service, so any supported deployment that exposes its API at the address in SPLASH_URL can be used.

Which Splash version enables POST arguments?

Splash 1.8 or newer adds the http_method and body arguments; with /execute, your Lua script must pass them to splash:go.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I do when a site needs multiple browser windows?

Use a modern headless-browser automation stack instead of Splash; Scrapy’s dynamic-content guidance identifies multiple-window interaction as a case where Splash may not be sufficient.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.