Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single public list that contains every URL on a domain. Build the most reliable inventory by merging the site’s XML sitemaps, an authenticated internal-link crawl, Google Search Console’s known and submitted URL data, URL Inspection checks, and a limited site: search. Keep the original URL, final response, crawlability and indexing status as separate fields; each source answers a different question.

First define what “all URLs” means

A domain can have URLs that are declared in a sitemap, linked from another page, discovered in JavaScript, known to Google, or currently indexed. Those sets overlap, but none is complete by itself. A useful audit therefore labels each record instead of treating every URL as equally verified.

  • Discovered: found in a sitemap, link, feed, script, Search Console, log, or another source.
  • Crawlable: your crawler can fetch it under the stated authentication, robots rules and crawl limits.
  • Indexed/servable: Google reports that it can index or serve the URL. This is different from merely discovering it.
  • Orphan: found by a sitemap, Search Console, logs or another source but not by the links your crawler followed.

Record the exact scheme and host you audited, including whether www, non-www, HTTP and HTTPS were treated as separate variants. Add the audit date, authentication state, user agent and crawl rules to the export.

The seven-source workflow

1. Read robots.txt at the exact host

Request https://example.com/robots.txt (substitute the scheme and host you are auditing). Save every User-agent, Allow, Disallow and fully qualified Sitemap: line. A sitemap URL in robots.txt must be a fully qualified URL. Do not assume that a rule for one host applies to another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt is a crawler instruction file, not an inventory. A disallowed path can still be known to Google, and a permitted path is not necessarily indexed.

2. Expand every sitemap and sitemap index

Download each sitemap named in robots.txt, plus the site’s advertised sitemap location if you have one. Follow nested sitemap indexes until you reach URL sets. Preserve the URL exactly as published, then store a normalized comparison form for deduplication.

  • Keep both the original URL and its final HTTP response after redirects.
  • Record lastmod when present, but do not treat it as proof that a page changed or was indexed.
  • Normalize host and case variants only for comparison; never overwrite the source value.
  • Save HTTP status, content type, fetch time and any parsing error.

Sitemaps help search engines discover URLs but do not guarantee that every listed item will be crawled or indexed.

3. Crawl internal links while authenticated when appropriate

Start with the canonical host and follow HTML links that remain inside the audited scope. If the site has members-only or staging areas that you are authorized to inspect, crawl with the appropriate session. Capture links from:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • HTML anchors, canonical elements and pagination.
  • XML or RSS feeds and media references.
  • Rendered JavaScript routes when a browser is required.
  • Links revealed after normal interactions, such as “load more” controls.

For every fetched URL, export status code, content type, canonical URL, noindex state, referring URL and crawl depth. Rate-limit requests, honor applicable robots rules and stop on a defined URL or time budget. A crawl is an observation of what your crawler could reach, not a claim that the domain contains nothing else.

4. Compare Search Console’s URL populations

In Google Search Console’s Page Indexing report, compare All known pages, All submitted pages and Unsubmitted pages only. The report’s example URL list is limited to 1,000 items, so it is useful for diagnosis and sampling, not a complete export of a large site.

Export what the interface makes available and mark the source and date. “Known” means Google has encountered the URL; it does not mean the URL is indexed or currently eligible to appear.

5. Use URL Inspection for disagreements

Inspect URLs that appear in one dataset but not another, or whose canonical, robots and indexing signals conflict. Check Discovery details, sitemap association, crawl and indexing status, rendered resources and blocking information. URL Inspection is a targeted diagnostic tool; it is not a bulk replacement for your inventory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Run a site: search as a spot check

Use site:example.com and narrower path variants to see what Google currently serves for the domain. Try important subdirectories and host variants. Treat result counts and visible results as an indexed sample, never as a complete URL count.

7. Reconcile and classify

Merge records by a carefully normalized URL while retaining every source flag. A practical classification set is:

Class Meaning Typical action
Sitemap-only Declared in a sitemap but not reached by your crawl Check links, authentication, redirects and orphan status
Crawl-only Reached through links or rendered routes but absent from sitemaps Decide whether it belongs in the sitemap or should be excluded
Search-Console-known Google has discovered it Inspect discovery and indexing diagnostics
Indexed/servable Evidence indicates Google can serve it Record the date and evidence source
Blocked Robots, authentication, network or application rules prevented a fetch Separate “not fetched” from “does not exist”
Redirected The requested URL resolves to another URL Keep both requested and final URLs
Duplicate Multiple URLs resolve to equivalent content or the same canonical Review canonical, redirect and parameter policy
Orphan Discovered by a non-link source but not by the link crawl Verify whether it is intentional and reachable

How the discovery methods differ

Method Coverage it provides Access Freshness What it cannot prove
Robots.txt and sitemaps URLs the owner declares Usually public Depends on publishing and refresh Crawl or indexing
Authenticated crawler URLs linked or rendered within its scope Public or authorized Current at crawl time Unlinked, blocked or undiscovered URLs
Search Console reports URLs Google knows or received in sitemaps Verified property Google’s reporting cycle A complete downloadable inventory
URL Inspection Detailed evidence for selected URLs Verified property Per inspection Bulk coverage
site: query Pages Google chooses to show Public search Search-result dependent Exact counts or exhaustive indexing

A small, repeatable Python crawler

The following standard-library script starts at a URL, follows same-host HTML links, records canonical and noindex signals, and writes a CSV. It is intentionally conservative: it does not bypass robots rules, execute JavaScript or log in. Use a browser-capable crawler for sites whose routes appear only after rendering.

import csv
import re
import time
from collections import deque
from html.parser import HTMLParser
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.request import Request, urlopen

START = "https://example.com/"
MAX_URLS = 5000
DELAY = 0.25

class LinkParser(HTMLParser):
    def __init__(self, base):
        super().__init__()
        self.base = base
        self.links = set()
        self.canonical = ""
        self.noindex = False
    def handle_starttag(self, tag, attrs):
        a = dict(attrs)
        if tag == "a" and a.get("href"):
            self.links.add(urljoin(self.base, a["href"]))
        if tag == "link" and a.get("rel", "").lower() == "canonical":
            self.canonical = urljoin(self.base, a.get("href", ""))
        if tag == "meta" and a.get("name", "").lower() == "robots":
            self.noindex = "noindex" in a.get("content", "").lower()

def same_host(url, host):
    p = urlparse(url)
    return p.scheme in ("http", "https") and p.netloc == host

start_host = urlparse(START).netloc
queue, seen, rows = deque([START]), set(), []
while queue and len(seen) < MAX_URLS:
    url = urldefrag(queue.popleft())[0]
    if url in seen or not same_host(url, start_host):
        continue
    seen.add(url)
    row = {"url": url, "status": "", "content_type": "", "canonical": "", "noindex": "", "error": ""}
    try:
        req = Request(url, headers={"User-Agent": "URLInventoryBot/1.0"})
        with urlopen(req, timeout=20) as r:
            row["status"] = r.status
            row["content_type"] = r.headers.get_content_type()
            body = r.read(2_000_000)
        if row["content_type"] == "text/html":
            parser = LinkParser(url)
            parser.feed(body.decode("utf-8", errors="replace"))
            row["canonical"], row["noindex"] = parser.canonical, parser.noindex
            for link in parser.links:
                clean = urldefrag(link)[0]
                if same_host(clean, start_host) and clean not in seen:
                    queue.append(clean)
    except Exception as exc:
        row["error"] = type(exc).__name__ + ": " + str(exc)
    rows.append(row)
    time.sleep(DELAY)

with open("crawl.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=rows[0].keys() if rows else ["url"])
    writer.writeheader(); writer.writerows(rows)
print(f"Wrote {len(rows)} rows to crawl.csv")

Run it only against properties you are allowed to crawl. Add sitemap URLs, redirects, response headers, feeds and rendered routes as separate inputs rather than pretending this script covers them.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

When your goal is to capture visual evidence of discovered pages rather than build the inventory itself, ScreenshotNeo provides a single-request website screenshot API. It can remove cookie/consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads and cache hits are not billed. Its MCP server gives Claude, Cursor and other MCP clients take_screenshot, get_page_info and capture_pdf tools.

For the full parameter list, see the ScreenshotNeo API documentation. Replace the target URL below with one from your inventory:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

You can also request full pages, selected elements, device presets, dark mode, custom CSS or JavaScript, waits, blocking rules, authentication headers, cookies, PDFs, signed links, asynchronous webhooks and bulk capture of up to 100 URLs per call. Every response includes X-Page-Verdict and X-Billed headers so you can see whether a clean shot was billed. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common gaps

The sitemap contains fewer URLs than the crawl

Those extra URLs are crawl-only. Check whether they are intentionally excluded, parameter variants, pagination, media, or pages that should be added to the sitemap. Confirm that redirects and canonical targets are represented correctly.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The sitemap URL returns an error or HTML

Verify the fully qualified URL from robots.txt, follow sitemap-index nesting, check authentication and content type, and retain the failure in your audit log. A failed download is not evidence that the sitemap is empty.

Search Console shows a URL your crawler never reached

Inspect it. It may be an orphan, an old URL, a redirect, an authenticated route or a URL discovered through an external source. Compare the inspected canonical, crawl status and blocking details with your crawl record.

URLs differ only by parameters or case

Store the raw URL, then define a documented normalization policy for comparisons. Do not discard parameters until you know whether they change content, tracking, filtering or pagination.

JavaScript pages are missing

Use a rendering crawler, capture routes after interactions and compare the rendered link graph with the server HTML. Record that rendered discovery used a different method and date.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The crawl is slow or repeatedly fails

Lower concurrency, add backoff for transient failures, cache responses, set a maximum body size and split the work by host or path. Keep timeouts and error types in the export so “failed to fetch” is not confused with “URL does not exist.”

Performance, reliability and cost controls

  • Start with cheap declarations: robots.txt and sitemaps reduce the initial crawl queue.
  • Use a bounded crawl: set maximum URLs, depth, response size, concurrency and elapsed time.
  • Preserve evidence: save response status, headers, canonical, noindex, source and timestamp for each record.
  • Repeat incrementally: recrawl changed paths and compare normalized URL sets instead of rebuilding everything on every run.
  • Separate environments: keep production, staging and authenticated hosts in distinct inventories.
  • Budget rendering: browser rendering and screenshot capture cost more time and resources than fetching static HTML, so reserve them for JavaScript-dependent routes and visual checks.

A final URL count should always state what it counts: sitemap entries, unique discovered URLs, crawlable responses, Search Console-known URLs or indexed samples. Those numbers are not interchangeable.

FAQ

Can I export every URL Google knows in one file?

Not from the standard Page Indexing example list: Google documents a 1,000-URL limit for that interface view. Combine available Search Console data with sitemaps, crawling and targeted inspection instead.

Should an orphan URL always be removed?

No. First determine whether it is intentionally private, legacy, campaign-specific or required by users. Then choose a link, redirect, noindex or removal action that matches its purpose.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should I report the result to another team?

Deliver the URL export with source flags, final responses, canonical and noindex fields, crawl date, host scope, authentication state and the rules used. That context makes the count reproducible.

Frequently Asked Questions

Can I export every URL Google knows in one file?

Not from the standard Page Indexing example list: Google documents a 1,000-URL limit for that interface view. Combine available Search Console data with sitemaps, crawling and targeted inspection instead.

Should an orphan URL always be removed?

No. First determine whether it is intentionally private, legacy, campaign-specific or required by users. Then choose a link, redirect, noindex or removal action that matches its purpose.

How should I report the result to another team?

Deliver the URL export with source flags, final responses, canonical and noindex fields, crawl date, host scope, authentication state and the rules used. That context makes the count reproducible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

For a defensible domain inventory, merge sitemap declarations, authenticated crawl results and Search Console evidence, then classify every URL by discovery, crawlability and indexing status. No single source can prove that it found every URL.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.