What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
To crawl many websites asynchronously, separate crawl orchestration from HTTP transport, keep a durable and deduplicated URL frontier, and limit work both globally and per domain. Use a crawl framework such as Scrapy when you want scheduling, retries, and crawl controls built in; use aiohttp when you need lower-level control and are prepared to build those controls yourself. There is no safe universal concurrency number: set limits from each site’s policy and your observed response behavior, not from a pages-per-second target.
What “asynchronous at scale” means
Async I/O lets a process make progress on other requests while waiting for network responses. It is useful when crawling spends substantial time waiting on DNS, connection setup, or remote servers. It does not remove the need to schedule URLs, deduplicate them, respect site policies, bound memory, or store results reliably. Nor does it make CPU-heavy parsing faster by itself.
A production crawler is more than a loop that launches HTTP requests. Its main components are seed ingestion, URL normalization, a durable frontier, deduplication, per-host or per-domain politeness state, fetch workers, parsing and extraction, persistence, and operational metrics. Keeping these concerns explicit helps a slow or failing domain stop affecting unrelated work.
Keep orchestration separate from transport
Scrapy is a crawl framework: it provides a scheduler and downloader, crawl-level concurrency and delay settings, retry behavior, and parsing and export facilities. Its asyncio integrations include AsyncCrawlerProcess and AsyncCrawlerRunner. aiohttp is a lower-level asynchronous HTTP client: it provides connection pooling and awaited response handling, but your application must supply the frontier, per-domain scheduling, retry policy, and other crawl orchestration.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Choose a framework or build the scheduler
| Approach | Best fit | What you must own |
|---|---|---|
| Scrapy-first | Structured crawls where built-in scheduling, retries, throttling, parsing, and feed export reduce application code. | Site-specific policy, crawl design, operational limits, and any multi-machine partitioning. |
| aiohttp-first | Applications that need direct control over asyncio and HTTP transport or already have a service architecture to integrate with. | Most crawl-level machinery: durable queues, deduplication, per-domain rate limits, retry budgets, and recovery. |
Use Scrapy if the work is primarily a crawl and its defaults and extension points fit. Choose aiohttp if transport control is the main requirement and you can implement and operate the scheduler around it. In either case, robots handling and resource limits need deliberate configuration.
Scrapy: configure broad-crawl behavior
For many domains, allow enough global concurrency to keep independent sites moving while keeping each domain slow enough to be polite. Scrapy’s default priority queue is optimized for a single domain; its documentation recommends DownloaderAwarePriorityQueue for broad crawls so downloader capacity across domains is considered.
# settings.py
ROBOTSTXT_OBEY = True
# Starting values, not universal safe limits. Adjust to policy and measurements.
CONCURRENT_REQUESTS = 32
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 1.0
# Useful for a broad crawl across many domains.
SCHEDULER_PRIORITY_QUEUE = "scrapy.pqueues.DownloaderAwarePriorityQueue"
# Keep retries bounded; retries consume crawl capacity.
RETRY_ENABLED = True
RETRY_TIMES = 2
# Use a finite timeout so stuck requests do not occupy capacity indefinitely.
DOWNLOAD_TIMEOUT = 30
These values are an illustrative starting configuration, not a recommended rate for every site. A domain’s published requirements and observed responses take precedence. Scrapy’s AutoThrottle can adjust per-site download delay based on load; it complements rather than replaces robots rules or explicit global limits.
aiohttp: reuse sessions and bound work
With aiohttp, reuse one ClientSession or a deliberately managed session pool. A session’s connector pools connections for reuse. The request call obtains response headers, while consuming the response body is a separate awaited operation; failing to read or release responses can undermine reuse and resource control.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
This small example demonstrates a bounded fetch queue and shared session for a fixed list of URLs. It is not a complete production crawler: it does not discover links, persist a frontier, implement robots policy, or apply per-host delays. Add those controls before expanding it to a real multi-domain crawl.
import asyncio
import aiohttp
URLS = [
"https://example.com/",
"https://www.iana.org/domains/reserved",
]
GLOBAL_LIMIT = 8
async def main():
semaphore = asyncio.Semaphore(GLOBAL_LIMIT)
timeout = aiohttp.ClientTimeout(total=30)
connector = aiohttp.TCPConnector(limit=GLOBAL_LIMIT)
async with aiohttp.ClientSession(
timeout=timeout,
connector=connector,
headers={"User-Agent": "ExampleResearchCrawler/1.0 (+contact: crawler@example.org)"},
) as session:
async def fetch(url):
async with semaphore:
async with session.get(url, allow_redirects=True) as response:
body = await response.read()
print(url, response.status, len(body), response.url)
await asyncio.gather(*(fetch(url) for url in URLS))
if __name__ == "__main__":
asyncio.run(main())
Install the dependency with python -m pip install aiohttp, save the code as fetch.py, and run python fetch.py. For production, replace the fixed list and gather fan-out with a bounded queue and workers. A large unbounded collection of tasks can consume memory before the network becomes the bottleneck.
How to set concurrency without overwhelming sites
Bound concurrency at two levels: total active requests across the crawl and simultaneous requests to an individual domain or host. A global limit protects your own file descriptors, sockets, memory, DNS, and downstream storage. A per-domain limit and delay protect each target and reduce the chance that it throttles or blocks the crawler.
- Start with site policy. Read each site’s robots.txt and any published crawl guidance. Translate applicable
Crawl-delayandRequest-ratedirectives into scheduler limits; Scrapy does not apply those directives automatically. - Set conservative per-domain limits. Begin slowly, then observe latency, status codes, and errors. Increase only when policy permits and the site remains responsive.
- Set a global ceiling. Choose a cap that your DNS, network sockets, CPU, memory, and storage can sustain. Global concurrency can be raised as the number of independent domains grows, but only while those resources remain healthy.
- Use backpressure. Bound queue size and active work so a fast discovery stage cannot overwhelm fetchers, parsers, or persistence.
- Reassess on failure. Rising timeouts, throttling, or server errors are signals to slow down, not to add more workers. Retries use capacity too.
More concurrency is not automatically more throughput. Once a target’s tolerance or your own bottleneck is reached, extra requests can increase errors, bans, retry traffic, and latency while reducing successful pages per unit of time. The reviewed primary guidance does not establish a universal pages-per-second figure: performance depends on target tolerance, response size and latency, DNS, parsing, storage, and retry behavior. Benchmark representative domains with explicit safety limits.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Respect robots.txt as a scheduler input
Robots Exclusion Protocol rules are crawler instructions, not authentication or a security boundary. RFC 9309, published by the IETF in September 2022, says: “These rules are not a form of access authorization.” A robots file neither grants permission to access protected material nor substitutes for access controls.
For a successful robots.txt download, follow the parseable rules that apply to the crawler’s user-agent. Match the most specific applicable rule rather than treating the file as a simple global allow or deny. Make robots retrieval a prerequisite for crawling that host, and record when the policy was fetched so workers do not independently fetch it for every URL.
Redirects, errors, and caching
Handle robots redirects and status failures deliberately. RFC 9309 distinguishes an unavailable file from an unreachable one: an unavailable response, such as a 4xx, may allow access to resources, while a server or network failure makes the robots file unreachable and calls for assuming complete disallow. Follow robots redirects according to the RFC, and apply the resulting rules to the original authority. Do not turn a temporary retrieval failure into a permanent allow rule.
Cache robots policy conservatively and refresh it according to the standard and your operational needs. Store the policy version and fetch time, and make the failure state explicit. This avoids silently crawling on stale assumptions after a robots fetch fails.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Build a durable frontier and prevent duplicate work
The frontier is the authoritative queue of URLs awaiting fetch. Normalize URLs consistently before admission: remove fragments, handle scheme and host casing consistently, and decide how your crawl treats query parameters, trailing slashes, and known tracking parameters. Over-aggressive normalization can merge pages that differ in content; insufficient normalization can cause duplicate fetches.
- Deduplicate durably. Record normalized URLs in durable storage rather than relying only on process memory. For concurrent workers, make admission atomic so two workers cannot both enqueue the same URL.
- Checkpoint progress. Persist pending work, completion state, retry count, and enough crawl metadata to resume after a worker or machine restarts.
- Track politeness state. Maintain next-eligible request time and in-flight count by host or domain. A domain waiting for its delay should not block eligible work for other domains.
- Bound retries. Use a finite retry budget and distinguish retryable transport failures from responses that should not be retried. Scrapy warns that slow or failing response retries can substantially reduce crawl capacity.
- Propagate cancellation. On shutdown, stop admitting new work, cancel or finish in-flight requests within a defined grace period, and preserve queue state for recovery.
Prefer APIs, bulk exports, search endpoints, or sitemaps when they can provide the needed data more directly than page crawling. This can reduce load on the site and simplify extraction.
Distribute a crawl across machines
Scrapy does not provide built-in multi-server distribution for one spider. A documented approach is to partition URL inputs and run those partitions on separate Scrapyd servers. For a broader distributed design, assign ownership of frontier partitions to workers or use a shared queue with durable, atomic claims.
Partitioning choices
- Partition by seed or URL range: simple to launch and recover, but URLs discovered across partitions can still collide.
- Partition by host: keeps a host’s politeness state local to one worker, which reduces coordination, but requires reassignment when workers change.
- Shared frontier: workers claim work from common durable storage, which eases balancing but requires atomic claims, distributed deduplication, and coordinated per-domain limits.
Whichever strategy you choose, specify who owns deduplication, retries, and the per-domain rate state. If two machines independently think they own a host, their combined request rate can exceed the intended limit. Checkpointed state and an explicit reassignment procedure are essential for crash recovery.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Operate for throughput, memory, and reliability
Scale domain parallelism only while the entire pipeline can keep up. Monitor queue depth, active requests, per-domain latency, HTTP status codes, retries, bytes received, parser lag, duplicate rate, DNS behavior, and storage latency. A growing frontier paired with a healthy fetch rate may indicate that parsing or persistence, rather than networking, is now the bottleneck.
- DNS and sockets: Check resolver latency, connection limits, and file descriptors before raising concurrency.
- Memory: Use bounded queues and disk-backed job state when the frontier or responses are large. Scheduling breadth-first rather than depth-first can change how much pending work accumulates.
- Response handling: Set timeouts and response-size limits appropriate to the job; stream large bodies when practical instead of retaining many full pages at once.
- Retries and timeouts: Reduce unnecessary retries and use finite timeouts for stuck requests. Repeated retries of slow targets lower capacity for useful work.
- Development controls: Disable cookies unless the crawl needs them, and use HTTP caching during development to avoid repeatedly fetching unchanged pages.
For repeatable performance checks, use a representative mix of domains and realistic page sizes, parsing, and persistence. Keep the safety limits in place during benchmarks; a result obtained by ignoring target-site tolerance is not a production capacity estimate.
Troubleshooting common crawl failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Many 429 or other throttling responses | Per-domain rate is too high, or the site has a policy the crawler has not applied. | Check robots.txt and site guidance, reduce that domain’s concurrency, increase its delay, and avoid aggressive retries. |
| Throughput falls as retries rise | Slow or failing hosts are consuming worker slots and retry capacity. | Bound retries, use finite timeouts, and isolate failing domains so they do not hold up other work. |
| Memory grows continuously | Unbounded task creation or frontier growth, or too many response bodies retained at once. | Use a bounded queue, cap in-flight work, persist frontier state, and stream or release response bodies promptly. |
| Same URL is fetched more than once | Normalization is inconsistent, or deduplication exists only in process memory or is not atomic across workers. | Canonicalize before queue admission and use a durable shared deduplication mechanism. |
| One host dominates or stalls a crawl | Scheduling is not downloader-aware, or waits and retries for that host block shared workers. | For broad Scrapy crawls, use the downloader-aware priority queue; isolate host state and let other domains proceed. |
| Distributed workers exceed a site’s intended rate | Each worker enforces a local limit without coordinating host ownership. | Assign each host to one worker or coordinate per-domain in-flight counts and next-eligible times centrally. |
| Robots policy seems inconsistent | Redirects, status failures, or stale cached policy are being treated as ordinary allow/deny responses. | Apply RFC 9309’s redirect and unavailable-versus-unreachable behavior, record fetch time and state, and do not convert network failure into an allow rule. |
When the job is screenshots rather than crawling
A crawler retrieves pages and manages a frontier; a screenshot API captures a rendered page. If the requirement is to collect visual snapshots rather than discover and fetch pages at scale, ScreenshotNeo is a separate tool for website screenshots and PDFs, not a crawler. Its API can be useful when you need a rendered capture for a URL.
Or skip the browser setup
One GET request returns the capture. Replace the URL with the page you need; use your API key from your account. The response is an image or PDF according to the requested options. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed before capture; supported cleanup also covers newsletter popups and chat widgets, and each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. The same features are available on every plan.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
FAQ
Does asyncio make HTML parsing faster?
No. Async I/O helps keep network work moving while requests wait. CPU-heavy parsing can still bottleneck the crawl; measure parser lag and scale or separate parsing only when measurements show it is needed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

