Free tools Windows power users keep installed
One-click scans. No signup required.
Scrapy Splash is a two-part system: scrapy-splash is the Scrapy integration, while Splash is a separate HTTP service that renders pages in a WebKit browser. Install the Python package, run a Splash server (Docker is the usual route), configure the documented middleware and request fingerprinter, then choose a rendering endpoint. Use render.html or render.json for straightforward pages; use /execute or /run when Lua must control navigation, JavaScript, cookies, or returned data.
This guide covers a working setup, Lua scripts, sessions, POST requests, compatibility limits, troubleshooting, and an alternative when maintaining a browser-rendering service is more work than you want.
How the Scrapy Splash architecture works
A normal Scrapy request downloads the response directly. A Splash request sends the target URL and rendering arguments to the Splash HTTP server. Splash loads the page in its WebKit engine, executes JavaScript, and sends the rendered result back to Scrapy.
- Scrapy: your spider, scheduling, parsing, retries, pipelines, and item extraction.
- scrapy-splash: Scrapy request classes, middleware, cookie handling, argument de-duplication, and request fingerprinting.
- Splash: the separately deployed rendering service and its HTTP API.
Installing scrapy-splash alone does not install or start Splash. Your Scrapy process must be able to reach the server address configured in SPLASH_URL.
#1 Best Overall
Prerequisites and installation
Use a supported Python environment
Current Scrapy installation guidance requires Python 3.10 or newer (CPython or PyPy) and recommends a dedicated virtual environment. Create one before installing project dependencies:
python -m venv .venv
# Linux/macOS
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
python -m pip install --upgrade pip
python -m pip install scrapy scrapy-splash
Run Splash with Docker
The commonly documented container command publishes Splash on port 8050:
docker run -p 8050:8050 scrapinghub/splash
Leave that process running. From the host, the service address is usually http://127.0.0.1:8050. If Scrapy runs in another container, do not use 127.0.0.1 to refer to the Splash container; use the Docker service name or another address reachable from the Scrapy container.
Configure Scrapy settings
Set the Splash URL and add the middleware in the documented order. The compression middleware must have a lower priority number than the Splash middleware so Splash responses are handled correctly:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
# settings.py
SPLASH_URL = 'http://127.0.0.1:8050'
DOWNLOADER_MIDDLEWARES = {
'scrapy_splash.SplashCookiesMiddleware': 723,
'scrapy_splash.SplashMiddleware': 725,
'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}
SPIDER_MIDDLEWARES = {
'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}
REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter'
SplashDeduplicateArgsMiddleware avoids repeatedly transferring identical large arguments. SplashRequestFingerprinter makes Splash arguments part of request identity, preventing Scrapy from treating materially different render requests as duplicates.
Choose the right Splash endpoint
render.html and render.json
Use these endpoints when you need a rendered document with little custom control. They accept normal Splash arguments such as a URL, viewport, wait time, or JavaScript settings. render.html returns HTML; render.json wraps rendered output and metadata in JSON.
/execute and /run
The Splash API describes execute and run as its most versatile endpoints because they execute arbitrary Lua rendering scripts. Choose them when you need to click, wait for a condition, evaluate JavaScript, manage cookies, make several navigations, or return a custom table rather than a standard response.
In scrapy-splash, an execution request normally uses endpoint='execute' and passes the script as lua_source. The /run endpoint is also Lua-based; use the endpoint that matches the response format and control flow required by your project.
A minimal Scrapy spider using Splash
This spider renders a page, waits briefly for client-side content, and parses the resulting HTML:
import scrapy
from scrapy_splash import SplashRequest
class ProductSpider(scrapy.Spider):
name = 'products'
allowed_domains = ['example.com']
start_urls = ['https://example.com/products']
def start_requests(self):
for url in self.start_urls:
yield SplashRequest(
url,
self.parse,
endpoint='render.html',
args={
'wait': 2,
'timeout': 90,
},
)
def parse(self, response):
for title in response.css('h2.product-title::text').getall():
yield {'title': title.strip()}
The timeout argument is a Splash-side limit for rendering. Scrapy’s own download timeout and retry settings still apply to the HTTP request made to Splash, so configure both layers for slow sites.
Writing Lua for /execute
The basic pattern
A Lua script defines main(splash). Navigate with splash:go, wait or evaluate JavaScript as needed, and return a string, value, or table:
script = r'''
function main(splash)
assert(splash:go(splash.args.url))
splash:wait(2)
return splash:evaljs("document.title")
end
'''
yield SplashRequest(
'https://example.com',
self.parse_title,
endpoint='execute',
args={'lua_source': script},
)
def parse_title(self, response):
# An execute response containing a scalar is available as response.text.
self.logger.info('Rendered value: %s', response.text)
splash.args.url receives the URL supplied by the Scrapy request. assert turns a failed navigation into a Lua traceback that you can inspect in Scrapy logs.
Recommended Free Tools
Return HTML and a value together
Returning a Lua table is useful when the spider needs both the rendered document and a computed value:
function main(splash)
assert(splash:go(splash.args.url))
splash:wait(1)
return {
title = splash:evaljs("document.title"),
html = splash:html()
}
end
Parse the JSON response according to the endpoint’s response format, and validate that the keys you expect are present before extracting data.
Waiting and JavaScript
A fixed splash:wait(seconds) is simple but can waste time or finish too early. For a page whose content appears after a known script action, call splash:evaljs and then wait for the resulting network and DOM work. If the site needs a browser feature WebKit does not implement, no amount of waiting will make that feature compatible.
Cookies, sessions, and request state
Splash is stateless per request. A later request does not automatically inherit cookies from an earlier one. To preserve a login or shopping session, pass cookies into Lua, initialize them before navigation, and return the updated cookie jar:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →function main(splash)
splash:init_cookies(splash.args.cookies)
assert(splash:go(splash.args.url))
return {
cookies = splash:get_cookies(),
html = splash:html()
}
end
On the Scrapy side, use the session support provided by scrapy-splash, including a consistent session_id, so related requests use the same logical cookie flow. Treat session identifiers as isolated per user or crawl; sharing one identifier across unrelated jobs can mix authentication state.
POST requests and cached Lua arguments
POST support
Splash 1.8 or newer is required for the http_method and body POST arguments. With /execute, the Lua script must pass those values to splash:go rather than assuming a GET:
function main(splash)
assert(splash:go{
url = splash.args.url,
http_method = splash.args.http_method,
body = splash.args.body,
})
return splash:html()
end
Set the method and body in the Scrapy request arguments, and ensure the target site also receives any required content type or authentication headers.
Cached arguments
Splash 2.1 or newer supports server-side caching of large static arguments such as lua_source. This can reduce request traffic and duplicate disk-queue data when many requests reuse the same script. It does not cache the target page’s changing content; it caches request arguments handled by Splash.
Compatibility: what can fail and why
Scrapy and Python
Use Python 3.10 or newer for current Scrapy installations, and verify that the scrapy-splash version you select supports the Scrapy release in your environment. Scrapy’s compatibility policy calls out backward incompatibilities in release notes and generally retains deprecated features for at least one year, but that policy does not guarantee that an unmaintained integration will track every new Scrapy release.
Rank #4
- Country of Origin:US
- CPSIA:N
- Hazardous?:No
- Tariff:4901990050
Splash and WebKit
The major compatibility boundary is the browser engine. Splash uses WebKit, and its FAQ warns that target sites can be incompatible with that engine. Modern sites may require browser APIs, JavaScript syntax, anti-bot behavior, or multi-window interaction that WebKit cannot provide.
Scrapy’s dynamic-content guidance positions Splash for JavaScript-rendered pages, while a modern headless browser may be necessary for on-the-fly DOM interaction or multiple windows. Choose Splash when its lightweight HTTP/Lua model fits the site; choose a current browser automation stack when the site’s behavior depends on newer browser capabilities.
Troubleshooting checklist
Connection refused or timeout before rendering
- Confirm the container is running with
docker ps. - Open the Splash address from the same network namespace as Scrapy.
- Check
SPLASH_URL; container-to-container setups usually require a service name, not127.0.0.1. - Raise Scrapy and Splash timeouts only after confirming the target page is reachable.
Duplicate requests or unexpected filtering
Ensure SplashDeduplicateArgsMiddleware and REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter' are present. Without the Splash fingerprinter, requests that differ only in rendering arguments can collide with ordinary Scrapy fingerprints.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Compressed or malformed responses
Check middleware priorities against the documented values. SplashMiddleware must run before HttpCompressionMiddleware in the effective order shown in the setup block.
Lua traceback or missing HTML
- Inspect the complete request, endpoint, and Lua traceback.
- Verify that
splash:gois wrapped inassertso navigation failures are visible. - Confirm every value used in Lua is supplied in
args. - For POSTs, verify Splash is 1.8 or newer and that the Lua call passes the method and body.
JavaScript content never appears
Increase the wait only if the page is genuinely slow. Then test whether the site depends on APIs unsupported by Splash’s WebKit engine. A modern headless browser may be required for complex interaction, newer browser features, or multiple windows.
Login state disappears
Remember that Splash requests are stateless. Initialize incoming cookies with splash:init_cookies, return splash:get_cookies(), and keep a stable scrapy-splash session_id for the sequence that must share state.
Operational and cost considerations
Self-hosting Splash gives you control over deployment and request volume, but you operate the container, networking, logs, retries, and upgrades. Keep verbose logging available while diagnosing failures: run the container with -v2 and inspect the full request and traceback, then reduce verbosity for routine operation.
Best Value
- Suitable for all kinds of project works
- Acid and toxic free
- Designed for easy usage
Rendering is more expensive than downloading raw HTML because Splash runs a browser engine for each request. Reuse Lua scripts, enable cached arguments where supported, avoid unnecessary waits, and request only the endpoint output your parser needs. There are no authoritative performance benchmarks in the documented material, so size capacity with measurements from your own target sites rather than a published throughput number.
Or skip the browser setup
If you need screenshots or PDFs rather than a self-managed Splash rendering service, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled.
Only clean shots are billed. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A direct call looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free, and every feature is available on every plan. Sign up free for ScreenshotNeo and start with the 1,000 no-card screenshots.
Frequently Asked Questions
Can I run Splash without Docker?
Yes. Docker is the commonly documented deployment, but Splash is a separate HTTP service, so any supported deployment that exposes its API at the address in SPLASH_URL can be used.
Which Splash version enables POST arguments?
Splash 1.8 or newer adds the http_method and body arguments; with /execute, your Lua script must pass them to splash:go.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsWhat should I do when a site needs multiple browser windows?
Use a modern headless-browser automation stack instead of Splash; Scrapy’s dynamic-content guidance identifies multiple-window interaction as a case where Splash may not be sufficient.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

