Recommended Free Tools
Use two filters for a crawler that should follow HTML pages: reject obvious file types while discovering links, then check each fetched response’s Content-Type before parsing it as HTML. The first filter can avoid unnecessary requests; the second is based on the returned representation rather than its URL. Neither is infallible, so define what to do with missing or misleading headers instead of silently treating every response as HTML.
Choose where to filter
A crawler can decide a URL is out of scope before requesting it, or decide a response is out of scope after it arrives. These are complementary checks, not substitutes. A path ending in .pdf is a useful clue, but it does not prove the response is a PDF; an extensionless URL can return one. The HTTP Content-Type header describes the media type of the returned representation, but servers may omit or misstate it.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Web-Crawler | $18.99 | Buy on Amazon |
| 2 |
|
A Handbook of Migrating Parallel Web Crawler | $78.95 | Buy on Amazon |
| 3 |
|
Web crawler Standard Requirements | $88.99 | Buy on Amazon |
| 4 |
|
Smart Web Crawler - эффективный рекурсивный захватчик... | $22.00 | Buy on Amazon |
| 5 |
|
Smart Web Crawler - Collecteur de ressources récursif efficace pour le Web (French Edition) | $44.00 | Buy on Amazon |
| Method | When it helps | Limitation |
|---|---|---|
| URL extension denylist | Skips familiar file types during link discovery, before a request is made. | Can miss extensionless files and exclude HTML served from a misleading suffix. |
Response Content-Type |
Lets the crawler decide whether to send a fetched response to an HTML parser. | The request has already happened; the header can be absent or wrong. |
HTTP HEAD preflight |
Can retrieve metadata without asking the server to send a response body. | Adds a request, and returned metadata may differ from the later GET. |
robots.txt |
Communicates crawl restrictions and helps manage crawler traffic. | Does not classify media types or guarantee that a URL is absent from search results. |
For most HTML-focused crawls, start with link filtering, fetch remaining URLs under the crawler’s normal limits, and gate HTML parsing on the response. Use HEAD only when the possible bandwidth savings justify an additional request and you have a fallback for unsupported or ambiguous results.
Filter discovered links in Scrapy
Scrapy’s LinkExtractor accepts deny_extensions. If you omit it, Scrapy uses its built-in IGNORED_EXTENSIONS list. That default is convenient, but inspect the list for your installed Scrapy version and your crawl’s purpose: a crawler intended to extract PDF text should not reject PDFs merely because they are not HTML.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
- SUPERHERO AND VEHICLE FIGURE SET: Many adventures with this Spidey and His Amazing Friends set, which includes a figure, vehicle, and accessory
- ARTICULATED FIGURE: This 4" figure features multiple points of articulation for lots of action
- TEAM SPIDEY ADVENTURES: Kids can be part of Team Spidey and create their own epic adventures with this Spidey and His Amazing Friends Vehicle Set
- INSPIRED BY MARVEL'S CHILDREN'S DRAWING: Little kids can imagine saving the day with their favorite superheroes with this Spidey and His Amazing Friends toy, inspired by the cute kids show
- ENDLESS ADVENTURES WITH SPIDEY AND HIS AMAZING FRIENDS TOYS: Other Spidey and His Amazing Friends Toys Available (sold separately and subject to availability)
Use the built-in extension list
For a conventional HTML crawl, allow Scrapy’s default ignored extensions to apply:
from scrapy.linkextractors import LinkExtractor
link_extractor = LinkExtractor()
Use this extractor in the spider’s link-following rules or wherever your spider extracts links. It affects discovered links only; it does not prove that a response from an allowed URL is HTML.
Specify a crawl-specific denylist
To make the policy explicit, provide the extensions you do not want to follow. Include the dot in each extension. This illustrative list covers common document, image, archive, audio and video suffixes; tailor it rather than assuming every crawl should exclude every listed format.
Rank #2
from scrapy.linkextractors import LinkExtractor
link_extractor = LinkExtractor(
deny_extensions=[
"pdf", "doc", "docx", "xls", "xlsx", "ppt", "pptx",
"jpg", "jpeg", "png", "gif", "webp", "svg",
"zip", "gz", "rar", "mp3", "mp4", "mov"
]
)
For example, if PDF documents are in scope but images are not, leave PDF extensions out of the denylist. Extension filtering is a URL-string heuristic: query strings, unusual routes, and server behavior can make suffix-based rules imperfect.
Reject selected links with process_value
LinkExtractor also provides a process_value hook. It can reject an individual extracted link by returning None, or return a modified value when you deliberately need to normalize a link. Keep such logic narrow: a broad rule can discard valid pages as well as unwanted downloads.
from urllib.parse import urlsplit
from scrapy.linkextractors import LinkExtractor
blocked = {".pdf", ".png", ".jpg", ".zip"}
def keep_non_blocked(link):
path = urlsplit(link).path.lower()
if any(path.endswith(ext) for ext in blocked):
return None
return link
link_extractor = LinkExtractor(process_value=keep_non_blocked)
This example checks the path and therefore avoids treating a query parameter such as ?format=pdf as a file suffix. It remains a heuristic: if the origin serves a PDF at /download?id=123, the path does not reveal that fact.
Rank #3
Check the response before HTML parsing
When the actual response arrives, inspect its Content-Type before passing its body to an HTML parser. A media type such as text/html is a straightforward signal. Decide explicitly whether to accept XHTML, how to handle a missing header, and whether an unexpected type should be dropped, logged, retried, or sent to a separate document pipeline.
Scrapy callback example
This callback illustrates a strict policy: parse only a response whose declared media type is text/html or application/xhtml+xml; log and skip anything else, including responses with no declared type.
Free tools Windows power users keep installed
One-click scans. No signup required.
import logging
from email.message import Message
import scrapy
logger = logging.getLogger(__name__)
class HtmlOnlySpider(scrapy.Spider):
name = "html_only"
start_urls = ["https://example.com/"]
def parse(self, response):
raw_type = response.headers.get(b"Content-Type", b"")
content_type = raw_type.decode("latin-1", errors="replace")
# Parse parameters such as charset, leaving the media type itself.
message = Message()
message["content-type"] = content_type
media_type = (message.get_content_type() if content_type else "").lower()
if media_type not in {"text/html", "application/xhtml+xml"}:
logger.info(
"Skipping non-HTML response url=%s status=%s content_type=%r",
response.url, response.status, content_type
)
return
title = response.css("title::text").get()
yield {"url": response.url, "title": title}
for href in response.css("a::attr(href)").getall():
yield response.follow(href, callback=self.parse)
The example shows the decision point, not a full production spider: add your own allowed-domain rules, crawl limits, duplicate handling, and error policy. Scrapy’s response selectors operate on a response body; making the media-type decision before using HTML selectors prevents a PDF or image body from being mistaken for a page.
Choose a policy for missing or suspicious headers
- Strict: skip a response if its media type is absent or not on your allowlist. This keeps the HTML pipeline predictable but can miss HTML from misconfigured servers.
- Permissive fallback: for a missing header only, inspect a small, bounded prefix for plausible HTML and record that the decision was inferred. This can recover misconfigured sites, but content sniffing is not definitive.
- Separate handling: route known document types to a PDF or other parser, and keep them out of the HTML parser. This is appropriate when those resources matter to the crawl.
Do not treat a header’s absence as proof that the body is non-HTML, or a declared HTML type as proof that the body is valid HTML. The response header is useful metadata, not a guarantee about server correctness.
Should you send a HEAD request first?
HTTP HEAD is intended to behave like GET except that the server does not send the response body. RFC 9110 says a server should generally send the headers it would send for GET, while allowing some headers to be omitted. In practice, some servers handle HEAD inconsistently, and the metadata may not match a later GET.
A preflight also costs an extra round trip for every URL that proceeds to a download. It is most useful when responses are large and the origin’s HEAD behavior is dependable. If you choose it, specify recovery behavior:
Best Value
- Send
HEADonly within the same domain, rate, timeout, and robots policy as the eventual crawl. - If the method is unsupported, fails, or returns ambiguous metadata, either fall back to a normal
GETwith a bounded body policy or record and skip the URL according to your crawl’s goal. - After any
GET, inspect that response’s own headers. Do not assume the earlierHEADsettled the representation type.
For general crawling, link filtering followed by a normal bounded GET and response check is simpler. A HEAD check is an optimization to evaluate, not a reliable replacement for checking the response you actually process.
Keep robots.txt separate from media filtering
robots.txt is a crawl-policy mechanism, not a file-type detector. Follow applicable crawl restrictions independently of your extension and content-type rules. Google’s robots.txt guidance notes that a URL disallowed from crawling can still be indexed when other pages link to it; a crawler therefore should not treat a robots rule as proof that a resource is non-HTML or gone from search.
Common problems and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| PDFs still reach the spider | The PDF uses an extensionless URL, or your discovery filter does not cover that link. | Check the fetched response’s Content-Type before HTML parsing; use a separate document handler if PDFs are in scope. |
| A real HTML page is skipped | Its URL has a denied suffix, or the server returned a missing or incorrect media type. | Review the matching URL and response headers. Narrow the denylist or define a logged fallback policy for missing/misdeclared types. |
| An image or binary response causes selector/parser errors | The callback parsed a body before checking its type. | Move the media-type gate ahead of HTML selectors and parser code. |
| HEAD succeeds but GET behaves differently | The origin omitted or varied metadata for HEAD, or the representation changed between requests. |
Make the final decision from the actual GET response and use a fallback if preflight is ambiguous. |
| URLs blocked by robots.txt still appear in search | Crawl restrictions prevent fetching content; they do not guarantee search engines will never know a linked URL. | Keep crawl compliance separate from indexing or removal goals; robots rules are not a content-type filter. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a web crawler or a way to filter non-HTML URLs. If your adjacent task is capturing a visual snapshot of a page, its one-request API is an alternative to launching and configuring a browser yourself. See the ScreenshotNeo documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Before the capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response includes X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
For screenshot capture rather than crawling, visit ScreenshotNeo or sign up free for 1,000 screenshots a month, with no card.
Frequently Asked Questions
Does an extension denylist prove a URL is not HTML?
No. It matches URL text, not the representation returned by the server. Treat it as a discovery shortcut and make the content decision from the response when needed.
Can robots.txt tell me whether a URL is HTML?
No. It is a crawl-policy mechanism, not a media-type signal.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →

