Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To list pages declared by a site’s robots.txt, fetch the site’s top-level /robots.txt, collect every case-insensitive Sitemap: record, then recursively read each sitemap index until you reach URL sets and their <loc> entries. The resulting inventory is a record of URLs declared through sitemaps—not proof that a site’s every page is live, unique, crawlable, canonical, or indexed.

What “every page” means in this workflow

robots.txt does not normally enumerate web pages. It is a Robots Exclusion Protocol file for crawler access requests. Sitemap records are an additional discovery mechanism. RFC 9309 specifies that robots.txt is served at the host’s top-level /robots.txt path, uses UTF-8, and should be delivered as text/plain. Google’s guidance allows multiple Sitemap: fields, requires each value to be an absolute URL, and does not tie the field to a particular User-agent group. A sitemap URL can point to another host.

A useful extractor therefore reports a sitemap-declared URL inventory. It should retain where every URL came from and distinguish successful retrievals from parser and network failures.

Important limits

  • A sitemap helps search engines discover URLs but does not guarantee that all listed items will be crawled or indexed.
  • A URL blocked by robots.txt can still be indexed if other pages link to it.
  • Sitemaps can be stale, contain duplicate locations, omit pages, or list URLs that now return errors.
  • “Every page” means every URL exposed by the discovered sitemap files, not every route that exists in an application or database.

Extraction pipeline

  1. Build the origin URL and request https://example.com/robots.txt (or the site’s supported HTTPS/HTTP protocol).
  2. Decode the response as UTF-8, preserving comments and recording the HTTP status and content type.
  3. Read every Sitemap: line. Trim surrounding whitespace, compare the field name case-insensitively, and validate that the value is an absolute URL.
  4. Fetch each sitemap. Decompress gzip content when the URL or response indicates compression.
  5. If the XML root is a sitemap index, enqueue each child sitemap and continue recursively. If it is a URL set, emit each <loc>.
  6. Track visited sitemap URLs to prevent cycles, apply a configurable nesting limit, and de-duplicate exact repeated page URLs.
  7. Export URLs plus provenance: robots.txt URL, sitemap URL, retrieval time, HTTP status, and parser result or error.

Complete Python extractor

The following script uses only Python’s standard library. It follows redirects, accepts compressed responses, handles XML namespaces, records failures, and writes both a URL list and an audit log. It deliberately preserves the exact text of each <loc> value; normalization is left to a later, explicit policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
#!/usr/bin/env python3
import argparse, gzip, json, sys
from collections import deque
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.request import Request, urlopen
import xml.etree.ElementTree as ET

UA = "SitemapInventory/1.0"

def now():
    return datetime.now(timezone.utc).isoformat()

def get(url, timeout):
    req = Request(url, headers={"User-Agent": UA, "Accept-Encoding": "gzip"})
    with urlopen(req, timeout=timeout) as r:
        data = r.read()
        if r.headers.get("Content-Encoding", "").lower() == "gzip":
            data = gzip.decompress(data)
        return data, r.status, r.headers.get("Content-Type", "")

def absolute(value):
    p = urlparse(value)
    return value if p.scheme in ("http", "https") and p.netloc else None

def robots_sitemaps(text):
    found = []
    for raw in text.splitlines():
        line = raw.split("#", 1)[0].strip()
        if ":" not in line:
            continue
        name, value = line.split(":", 1)
        if name.strip().lower() != "sitemap":
            continue
        candidate = absolute(value.strip())
        if candidate and candidate not in found:
            found.append(candidate)
    return found

def local(root):
    return root.tag.rsplit("}", 1)[-1].lower()

def children(root, name):
    for node in root.iter():
        if node.tag.rsplit("}", 1)[-1].lower() == name:
            if node.text and node.text.strip():
                yield node.text.strip()

def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("site", help="https://example.com")
    ap.add_argument("--timeout", type=float, default=30)
    ap.add_argument("--max-depth", type=int, default=20)
    ap.add_argument("--out", default="sitemap-inventory.json")
    args = ap.parse_args()
    origin = args.site.rstrip("/")
    robots_url = origin + "/robots.txt"
    audit, errors, pages, seen_pages, seen_maps = [], [], [], set(), set()
    try:
        raw, status, content_type = get(robots_url, args.timeout)
        text = raw.decode("utf-8-sig")
        maps = robots_sitemaps(text)
        audit.append({"kind":"robots", "url":robots_url, "retrieved_at":now(), "status":status, "content_type":content_type})
    except Exception as e:
        print(f"robots.txt failed: {e}", file=sys.stderr); return 2
    queue = deque((u, 0) for u in maps)
    while queue:
        sitemap_url, depth = queue.popleft()
        if sitemap_url in seen_maps: continue
        seen_maps.add(sitemap_url)
        record = {"kind":"sitemap", "url":sitemap_url, "retrieved_at":now(), "depth":depth}
        try:
            raw, status, content_type = get(sitemap_url, args.timeout)
            record.update(status=status, content_type=content_type, parser_result="ok")
            root = ET.fromstring(raw)
            kind = local(root)
            if kind == "sitemapindex":
                for child in children(root, "loc"):
                    child_url = absolute(child)
                    if child_url and depth + 1 <= args.max_depth:
                        queue.append((child_url, depth + 1))
                    elif not child_url:
                        errors.append({"url":sitemap_url, "error":"non-absolute child loc", "value":child})
            elif kind == "urlset":
                for value in children(root, "loc"):
                    if value not in seen_pages:
                        seen_pages.add(value); pages.append(value)
            else:
                record["parser_result"] = "unknown_root"
                errors.append({"url":sitemap_url, "error":"XML root is not sitemapindex or urlset"})
        except Exception as e:
            record.update(parser_result="error", error=str(e))
            errors.append({"url":sitemap_url, "error":str(e)})
        audit.append(record)
    output = {"robots_txt":robots_url, "retrieved_at":now(), "sitemaps":audit,
              "urls":pages, "errors":errors}
    with open(args.out, "w", encoding="utf-8") as f:
        json.dump(output, f, ensure_ascii=False, indent=2)
    print(f"Wrote {len(pages)} URLs to {args.out}; {len(errors)} errors")

if __name__ == "__main__": main()

Run it

python3 sitemap_inventory.py https://example.com --out example-inventory.json

The JSON file’s sitemaps array shows each fetch and parser result; urls contains exact, de-duplicated locations; errors keeps failures instead of silently dropping them. A nonzero exit code means robots.txt itself could not be retrieved.

Handling robots.txt correctly

Protocol and redirects

Start at the site’s origin and request the top-level path, not /sitemap.xml guessed from convention. Follow HTTP redirects but retain the original robots.txt URL in your audit. If a site is available only over HTTP, use that supported protocol and record it.

Comments, case, and malformed lines

Comments can follow a record and should not become part of the URL. Field-name matching is commonly made case-insensitive, while the URL value should be trimmed but otherwise preserved. Ignore unrelated directives such as User-agent, Allow, and Disallow when collecting sitemap records; a Sitemap record is independent of those groups.

Absolute and cross-host locations

Reject relative sitemap values because the documented Sitemap field requires absolute URLs. Do not restrict accepted hosts to the robots.txt host: a sitemap may be hosted elsewhere. Apply your own security policy before fetching untrusted cross-host URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Traversing sitemap XML

Sitemap indexes

An index contains child sitemap locations. Use a queue or recursion, maintain a visited set, and impose a depth limit. The visited set protects against cycles and duplicate references. Preserve each child’s parent in your audit if you need a full lineage report.

URL sets

A URL set contains one or more <url> elements, each with a <loc>. XML namespaces are normal, so match the local element name rather than assuming one namespace URI. Emit the location exactly as supplied unless your specification explicitly permits canonicalization.

Compression and XML failures

Some sitemap files are gzip-compressed. Handle the HTTP Content-Encoding header and, if your implementation supports it, a .gz URL. Treat malformed XML, an unexpected root element, invalid UTF-8, and decompression failures as auditable errors. Continuing with other sitemap files is usually more useful than aborting the entire inventory.

Output design and quality checks

Field Recommended value Why it matters
robots_txt_url Original top-level URL Identifies the discovery source.
sitemap_url Every fetched sitemap or index Supports reproducibility and troubleshooting.
retrieved_at UTC timestamp per request Shows when the inventory was observed.
http_status Response status Separates successful and failed fetches.
parser_result urlset, sitemapindex, unknown_root, or error Makes parser behavior inspectable.
loc Exact XML text Avoids silently changing publisher data.

Checks after extraction

  • Count sitemap records and compare them with the number fetched.
  • Look for duplicate sitemap URLs and duplicate page locations.
  • Optionally issue separate HTTP checks to classify pages as live, redirected, missing, or blocked; do not confuse those results with sitemap extraction.
  • Flag non-HTTP(S) locations, malformed URLs, and unexpected hosts for review.
  • Keep the raw robots.txt and sitemap responses when you need an audit trail, subject to the site’s terms and your retention policy.

Performance, reliability, and safety

Control concurrency

Large indexes can reference many files. A bounded worker pool can reduce elapsed time, but use conservative concurrency, per-request timeouts, retries with backoff, and a clear maximum depth. Caching responses during one run prevents repeated downloads when the same sitemap is referenced more than once.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Separate transient and permanent errors

Retry timeouts and temporary server errors a limited number of times. Do not endlessly retry authentication failures, malformed XML, or consistently unsupported content. Record the final status and error text.

Protect the extractor

Cross-host sitemap URLs and redirects mean the extractor may contact servers outside the original site. Restrict schemes to HTTP and HTTPS, enforce response-size limits, reject dangerous private-network destinations where appropriate, and avoid expanding entities or external XML references. These controls matter when URLs come from untrusted input.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

robots.txt returns 404

There may be no declared sitemap. Confirm the origin, protocol, redirect chain, and exact top-level path. Do not silently substitute a guessed sitemap URL unless your product specification says to do so.

No Sitemap records are found

Check comments, whitespace, capitalization, and whether the response is actually a robots.txt file. A site can publish a sitemap without advertising it in robots.txt; discovering that requires a separate, explicitly documented strategy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

XML parses but yields zero URLs

You may have fetched a sitemap index and failed to enqueue its children, encountered namespaces your parser does not handle, or received an empty/stale file. Inspect the root element and raw response.

One child sitemap fails

Keep the successful files, record the failed URL, status, and parser error, and retry transient failures. The final inventory should state that it is partial.

Duplicate URLs appear

Exact de-duplication is safe for a basic inventory. Do not collapse case variants, fragments, encoded forms, or redirects unless you define a canonicalization policy and report it separately.

Requests time out or are blocked

Increase the timeout modestly, honor redirects, identify your client, and reduce concurrency. A bot check or access denial is a retrieval failure—not evidence that the sitemap has no URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Or skip the browser setup

If your next step is to capture each discovered page visually, ScreenshotNeo provides a GET-based screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the page verdict and billing result in headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

See the ScreenshotNeo documentation for all options. A one-call capture looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to try it.

Frequently Asked Questions

Can robots.txt list every URL on a website?

No. It can advertise sitemap files, and those files can enumerate declared locations, but pages may be omitted, stale, duplicated, or inaccessible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a sitemap index different from a sitemap URL set?

Yes. An index points to child sitemap files; a URL set contains page locations. An extractor must distinguish the XML root and recurse through indexes.

Should I obey Disallow while extracting sitemap URLs?

Keep robots directives in your audit, but do not confuse access rules with the Sitemap field. Whether your own crawler may fetch listed pages is a separate policy decision.

Why preserve provenance for each URL?

It lets you identify the sitemap, retrieval time, response status, and parser result that produced an entry, which is essential when inventories change or files fail.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.