Assuming this title refers to Crawl4AI—the project whose recent releases combine web scraping, PDF ingestion, Docker infrastructure and security work—the key change is that PDF downloads now have their own security boundary. Crawl4AI v0.9.4 is listed as the latest release dated 23 September 2026. Its v0.9.3 release added redirect, size, page-count, file-write and HTML-escaping controls for PDFs, while v0.9.0 made the self-hosted Docker API secure by default. These controls reduce risk, but untrusted-document processing still needs isolation, timeouts and resource limits.
What this release series is actually changing
The releases address two different paths:
- Document path: A Docker request could select
PDFContentScrapingStrategy. That path fetched a PDF with Pythonrequests, outside browser egress and resource controls. Version 0.9.3 places validation and limits at this separate trust boundary. - Server path: Version 0.9.0 changes the self-hosted Docker HTTP server’s default posture. Authentication is enabled by default, binding is loopback unless a token is configured, and network request bodies are treated as untrusted input.
The v0.9.0 notes say the core in-process Python library was not changed by the Docker hardening. A deployment that imports Crawl4AI directly therefore has a different exposure than the Docker API and must be secured at its own process and network boundaries.
Current version and security status
Crawl4AI’s release listing identifies v0.9.4, released 23 September 2026, as the latest version at that date. The project’s security overview reports that this release fixes two SSRF paths and an untrusted-configuration bypass. Specifically, robots.txt and link-preview fetching are routed through a pinning egress proxy, and nested typed objects are rechecked against the untrusted-configuration gate. These are project-reported fixes, not an independent penetration test or a guarantee for every deployment.
Check the release and migration documentation again when you deploy: version status, defaults and configuration names can change.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
What v0.9.3 fixes in the PDF path
Local image writes are no longer caller-controlled
An untrusted request body can no longer choose arbitrary values for save_images_locally and image_save_dir. The server filters those fields and forces extract_images off for untrusted bodies. This prevents a network caller from selecting an output directory on the server’s filesystem.
Redirects are checked at every hop
Checking only the original PDF URL is insufficient: a public URL can redirect to an internal address. Crawl4AI manually validates PDF redirects, limits the chain to five hops, and validates the peer IP of the response. That combination addresses both destination changes and the address actually connected to the server.
PDF size, page count and execution time are bounded
The v0.9.3 notes set max_pdf_bytes to 100 MiB and max_pdf_pages to 2,000 pages. Untrusted Docker request bodies cannot raise those values above the caps. The Docker configuration also changes limits.wall_clock_s to 300 seconds. These are project limits, not universal values for every workload; high-volume systems may need stricter limits at an outer queue or gateway.
Extracted text is escaped before HTML insertion
Paragraph text parsed from a PDF is escaped before it is inserted into cleaned_html. The release also removes a Playground viewer round-trip that interpreted crawled content as live HTML. Treat extracted text as data in downstream templates, databases and dashboards; do not assume that escaping in one output makes every later sink safe.
Rank #2
PDF strategy selection is wired correctly
When a Docker request selects PDFContentScrapingStrategy, the server routes it to PDFCrawlerStrategy automatically. That pairing works without a second manual strategy setting. It improves usability, but it does not remove the need to validate URLs, credentials and document policy before submitting a job.
Secure-by-default Docker behavior in v0.9.0
Authentication and network binding
The self-hosted Docker API enables authentication by default and binds to loopback unless a token is configured. Do not expose the container directly to a public interface while assuming the default is sufficient; confirm the effective bind address, token handling and reverse-proxy policy in the version you run.
Request bodies are an untrusted configuration channel
Every field supplied over the network should be treated as attacker-controlled. The server’s untrusted-configuration gate is important because a harmless-looking option can affect filesystem writes, network access, browser behavior or resource consumption. Apply the same principle to nested objects, not only top-level keys.
Artifacts are fetched through authenticated endpoints
Screenshot and PDF output moved to artifact identifiers retrieved through an authenticated endpoint, with a time-to-live and storage quota. This is a breaking change for the self-hosted HTTP server. Update clients that previously expected an output file to be returned directly, and verify artifact expiry and quota behavior under retry conditions.
Recommended Free Tools
A safe deployment procedure
- Pin the version. Record the exact Crawl4AI image or package version, rather than relying on a moving latest tag. For this release series, verify whether you intend v0.9.3 PDF behavior or the v0.9.4 security fixes.
- Keep the API private. Start with loopback or a private network, require the configured token, and place any public gateway in front of the service. Allow only the callers that need crawling.
- Separate browser and PDF egress policy. The PDF path uses a direct HTTP fetch, so browser routing rules do not automatically protect it. Apply outbound firewall or proxy rules to the container and validate every redirect destination.
- Set outer limits. Enforce queue concurrency, request timeout, response-body limit and job deadline outside the parser as well as the documented 100 MiB, 2,000-page and 300-second settings.
- Use a non-privileged runtime. Give the container a read-only root filesystem where practical, a writable directory only for controlled artifacts, and no host mounts that contain secrets or application data.
- Isolate document processing. Use a dedicated worker or sandbox for untrusted PDFs. Do not combine PDF parsing with credentials, production databases or administrative network access.
- Validate outputs. Store extracted text as text, sanitize at the final rendering context, and treat generated HTML, screenshots and PDFs as untrusted artifacts until scanned and access-controlled.
- Monitor and retire artifacts. Record URL, redirect chain, byte count, page count, verdict and elapsed time. Remove artifacts when their TTL expires and alert on repeated limit hits.
Why PDF files need layered defenses
Apache PDFBox summarizes the boundary plainly: “Processing untrusted PDFs is supported, but only to a defined extent.” Malformed documents can consume excessive CPU, memory, recursion depth or processing time. Its guidance recommends timeouts, memory limits, resource controls and sandboxing for applications processing untrusted documents at scale.
ASD system-hardening guidance adds host-level controls: use hardened PDF-application settings and prevent users from changing them. One listed control is blocking PDF applications from creating child processes. These safeguards reduce blast radius; none proves that a document is safe.
The UK Software Security Code of Practice is a voluntary baseline of 14 principles for software supplied to business customers. It is useful for documenting ownership, vulnerability handling and resilience expectations, but it is not a Crawl4AI certification.
Browser-mediated scraping versus direct PDF fetching
| Decision point | Browser-mediated page | Separate PDF request |
|---|---|---|
| Network controls | Browser egress and resource policies can apply. | Must enforce URL, redirect and peer validation in the PDF path. |
| Content limit | Page and resource limits depend on browser configuration. | Crawl4AI v0.9.3 caps PDFs at 100 MiB and 2,000 pages for untrusted Docker bodies. |
| Output risk | Rendered DOM can contain active HTML or scripts. | Parsed text must be escaped before insertion into cleaned_html and again at downstream sinks. |
| Failure mode | Browser checks, consent flows or timeouts can prevent a usable capture. | Redirect chains, parser exhaustion and oversized files require explicit handling. |
Choose the browser path when you need JavaScript-rendered content and can tolerate its runtime. Choose direct PDF fetching when the source is a document and you can enforce the stricter network and parser boundary. Do not assume that a setting on one path protects the other.
Rank #4
Generating screenshots or PDFs without operating a browser
If your pipeline only needs a rendered page image or PDF, a dedicated capture API can remove browser installation and maintenance from the worker. ScreenshotNeo is a website screenshot API and MCP server. It accepts one request for PNG, JPEG, WebP or PDF output, can load lazy images for full-page captures, and supports selectors, device presets, custom CSS and JavaScript, waits, headers, cookies, proxy-style blocking controls, geolocation, signed links and asynchronous webhooks.
Or skip the browser setup
Use the API documented at https://screenshotneo.com/docs/. The following calls are runnable after replacing YOUR_API_KEY and the target URL.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
The Free plan includes 1,000 screenshots each month with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Operational limits, reliability and cost
- Concurrency: PDF parsing is CPU- and memory-intensive. Bound simultaneous jobs and use a queue so one batch cannot starve interactive requests.
- Retries: Retry transient network failures with backoff, but do not blindly retry parser-limit failures, authentication errors or blocked destinations.
- Idempotency: Key jobs by source URL, document hash and requested options. This prevents duplicate processing when a client retries after an artifact timeout.
- Caching: Cache only when the source, authorization context and retention policy permit it. Never let a shared cache expose a private document.
- Cost accounting: Track bytes, pages, wall-clock time, worker memory and external API calls. A low request count can still be expensive when documents are large or repeatedly retried.
- Data residency: If documents leave your network, document the provider’s region and retention. Adobe says its server-side PDF Services and PDF Embed components run in Adobe Document Cloud on AWS infrastructure in US-East and EMEA, with a customer-selectable processing region; uploaded user-generated content is temporarily cached during normal operations.
Cloud and compliance considerations
AWS states, “Security is a shared responsibility between AWS and you.” A managed cloud provider protects underlying infrastructure, while you remain responsible for configuration, identities, data sensitivity and applicable requirements. Cloud hosting does not remove the need to restrict egress, protect credentials or delete artifacts.
For hosted PDF processing, inspect transport encryption, region, retention, access controls and document permissions. Adobe documents TLS 1.2 or greater in transit. It also notes that some PDF permission settings prevent API processing, and that password-protected PDFs require the password and authorization to remove protection. Do not send a restricted document until those conditions are understood.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF request is rejected after a redirect | A hop resolves to a disallowed address, or the five-hop limit is exceeded. | Inspect the complete redirect chain, allow only intended destinations and avoid URL shorteners. |
| Images are not extracted | The request body is untrusted and extract_images is forced off. |
Use a controlled, authenticated workflow and a permitted output directory; never re-enable arbitrary paths for public callers. |
| Large document stops near the limit | The 100 MiB or 2,000-page cap was reached. | Reject or pre-process the source, split it in a trusted worker, or apply a stricter business rule. |
| Job ends at five minutes | The Docker wall-clock default is 300 seconds. | Reduce document size or concurrency and review the version’s supported limit configuration; do not let callers raise it unchecked. |
| Client cannot download a prior artifact | v0.9.0 artifact identifiers use an authenticated endpoint with TTL and quota. | Refresh credentials, fetch before expiry and monitor storage quota; update clients for the breaking API behavior. |
| Rendered output contains unexpected markup | Downstream code treated extracted text as HTML. | Escape for the final context (HTML, JSON, SQL or shell) and keep content in a data field until rendering. |
| Screenshot API response is not billed | The page was a bot check, blank, timed out, failed to load or served from cache. | Read the X-Page-Verdict and X-Billed headers, then fix the target or retry only when appropriate. |
FAQ
Does this make Crawl4AI safe for every website and PDF?
No. The releases add specific controls and safer defaults. You still need permission to access a site, an isolation model, secret management and policies appropriate to your documents.
Are the 100 MiB, 2,000-page and 300-second values performance benchmarks?
No. They are project configuration limits and defaults documented for the v0.9.3 Docker PDF path, not independently measured throughput or capacity results.
Is the EDPB material a final legal rule for scraping?
No. The referenced page describes Guidelines 03/2026 as open for feedback through 30 October 2026. Treat it as a consultation notice, and obtain jurisdiction-specific legal advice for your use case.
Frequently Asked Questions
Which Crawl4AI component changed in v0.9.0?
The secure-by-default changes apply to the self-hosted Docker HTTP server. The release notes state that the core in-process Python library was unchanged.
What should I log for a PDF crawl?
At minimum, record the source URL, every redirect destination, connected peer, byte and page counts, elapsed time, limit failures, authentication result and artifact identifier.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

