Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteThere is no single public list that contains every URL on a domain. Build the most reliable inventory by merging the site’s XML sitemaps, an authenticated internal-link crawl, Google Search Console’s known and submitted URL data, URL Inspection checks, and a limited site: search. Keep the original URL, final response, crawlability and indexing status as separate fields; each source answers a different question.
First define what “all URLs” means
A domain can have URLs that are declared in a sitemap, linked from another page, discovered in JavaScript, known to Google, or currently indexed. Those sets overlap, but none is complete by itself. A useful audit therefore labels each record instead of treating every URL as equally verified.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The Proxy Playbook: The Complete Guide to Proxy Servers: How to Source, Test, and Scale Residential,... | $29.95 | Buy on Amazon |
| 2 |
|
How to Host your own Web Server | $15.60 | Buy on Amazon |
- Discovered: found in a sitemap, link, feed, script, Search Console, log, or another source.
- Crawlable: your crawler can fetch it under the stated authentication, robots rules and crawl limits.
- Indexed/servable: Google reports that it can index or serve the URL. This is different from merely discovering it.
- Orphan: found by a sitemap, Search Console, logs or another source but not by the links your crawler followed.
Record the exact scheme and host you audited, including whether www, non-www, HTTP and HTTPS were treated as separate variants. Add the audit date, authentication state, user agent and crawl rules to the export.
The seven-source workflow
1. Read robots.txt at the exact host
Request https://example.com/robots.txt (substitute the scheme and host you are auditing). Save every User-agent, Allow, Disallow and fully qualified Sitemap: line. A sitemap URL in robots.txt must be a fully qualified URL. Do not assume that a rule for one host applies to another.
#1 Best Overall
Robots.txt is a crawler instruction file, not an inventory. A disallowed path can still be known to Google, and a permitted path is not necessarily indexed.
2. Expand every sitemap and sitemap index
Download each sitemap named in robots.txt, plus the site’s advertised sitemap location if you have one. Follow nested sitemap indexes until you reach URL sets. Preserve the URL exactly as published, then store a normalized comparison form for deduplication.
- Keep both the original URL and its final HTTP response after redirects.
- Record
lastmodwhen present, but do not treat it as proof that a page changed or was indexed. - Normalize host and case variants only for comparison; never overwrite the source value.
- Save HTTP status, content type, fetch time and any parsing error.
Sitemaps help search engines discover URLs but do not guarantee that every listed item will be crawled or indexed.
3. Crawl internal links while authenticated when appropriate
Start with the canonical host and follow HTML links that remain inside the audited scope. If the site has members-only or staging areas that you are authorized to inspect, crawl with the appropriate session. Capture links from:
- HTML anchors, canonical elements and pagination.
- XML or RSS feeds and media references.
- Rendered JavaScript routes when a browser is required.
- Links revealed after normal interactions, such as “load more” controls.
For every fetched URL, export status code, content type, canonical URL, noindex state, referring URL and crawl depth. Rate-limit requests, honor applicable robots rules and stop on a defined URL or time budget. A crawl is an observation of what your crawler could reach, not a claim that the domain contains nothing else.
4. Compare Search Console’s URL populations
In Google Search Console’s Page Indexing report, compare All known pages, All submitted pages and Unsubmitted pages only. The report’s example URL list is limited to 1,000 items, so it is useful for diagnosis and sampling, not a complete export of a large site.
Export what the interface makes available and mark the source and date. “Known” means Google has encountered the URL; it does not mean the URL is indexed or currently eligible to appear.
5. Use URL Inspection for disagreements
Inspect URLs that appear in one dataset but not another, or whose canonical, robots and indexing signals conflict. Check Discovery details, sitemap association, crawl and indexing status, rendered resources and blocking information. URL Inspection is a targeted diagnostic tool; it is not a bulk replacement for your inventory.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute6. Run a site: search as a spot check
Use site:example.com and narrower path variants to see what Google currently serves for the domain. Try important subdirectories and host variants. Treat result counts and visible results as an indexed sample, never as a complete URL count.
7. Reconcile and classify
Merge records by a carefully normalized URL while retaining every source flag. A practical classification set is:
| Class | Meaning | Typical action |
|---|---|---|
| Sitemap-only | Declared in a sitemap but not reached by your crawl | Check links, authentication, redirects and orphan status |
| Crawl-only | Reached through links or rendered routes but absent from sitemaps | Decide whether it belongs in the sitemap or should be excluded |
| Search-Console-known | Google has discovered it | Inspect discovery and indexing diagnostics |
| Indexed/servable | Evidence indicates Google can serve it | Record the date and evidence source |
| Blocked | Robots, authentication, network or application rules prevented a fetch | Separate “not fetched” from “does not exist” |
| Redirected | The requested URL resolves to another URL | Keep both requested and final URLs |
| Duplicate | Multiple URLs resolve to equivalent content or the same canonical | Review canonical, redirect and parameter policy |
| Orphan | Discovered by a non-link source but not by the link crawl | Verify whether it is intentional and reachable |
How the discovery methods differ
| Method | Coverage it provides | Access | Freshness | What it cannot prove |
|---|---|---|---|---|
| Robots.txt and sitemaps | URLs the owner declares | Usually public | Depends on publishing and refresh | Crawl or indexing |
| Authenticated crawler | URLs linked or rendered within its scope | Public or authorized | Current at crawl time | Unlinked, blocked or undiscovered URLs |
| Search Console reports | URLs Google knows or received in sitemaps | Verified property | Google’s reporting cycle | A complete downloadable inventory |
| URL Inspection | Detailed evidence for selected URLs | Verified property | Per inspection | Bulk coverage |
site: query |
Pages Google chooses to show | Public search | Search-result dependent | Exact counts or exhaustive indexing |
A small, repeatable Python crawler
The following standard-library script starts at a URL, follows same-host HTML links, records canonical and noindex signals, and writes a CSV. It is intentionally conservative: it does not bypass robots rules, execute JavaScript or log in. Use a browser-capable crawler for sites whose routes appear only after rendering.
import csv
import re
import time
from collections import deque
from html.parser import HTMLParser
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.request import Request, urlopen
START = "https://example.com/"
MAX_URLS = 5000
DELAY = 0.25
class LinkParser(HTMLParser):
def __init__(self, base):
super().__init__()
self.base = base
self.links = set()
self.canonical = ""
self.noindex = False
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag == "a" and a.get("href"):
self.links.add(urljoin(self.base, a["href"]))
if tag == "link" and a.get("rel", "").lower() == "canonical":
self.canonical = urljoin(self.base, a.get("href", ""))
if tag == "meta" and a.get("name", "").lower() == "robots":
self.noindex = "noindex" in a.get("content", "").lower()
def same_host(url, host):
p = urlparse(url)
return p.scheme in ("http", "https") and p.netloc == host
start_host = urlparse(START).netloc
queue, seen, rows = deque([START]), set(), []
while queue and len(seen) < MAX_URLS:
url = urldefrag(queue.popleft())[0]
if url in seen or not same_host(url, start_host):
continue
seen.add(url)
row = {"url": url, "status": "", "content_type": "", "canonical": "", "noindex": "", "error": ""}
try:
req = Request(url, headers={"User-Agent": "URLInventoryBot/1.0"})
with urlopen(req, timeout=20) as r:
row["status"] = r.status
row["content_type"] = r.headers.get_content_type()
body = r.read(2_000_000)
if row["content_type"] == "text/html":
parser = LinkParser(url)
parser.feed(body.decode("utf-8", errors="replace"))
row["canonical"], row["noindex"] = parser.canonical, parser.noindex
for link in parser.links:
clean = urldefrag(link)[0]
if same_host(clean, start_host) and clean not in seen:
queue.append(clean)
except Exception as exc:
row["error"] = type(exc).__name__ + ": " + str(exc)
rows.append(row)
time.sleep(DELAY)
with open("crawl.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=rows[0].keys() if rows else ["url"])
writer.writeheader(); writer.writerows(rows)
print(f"Wrote {len(rows)} rows to crawl.csv")
Run it only against properties you are allowed to crawl. Add sitemap URLs, redirects, response headers, feeds and rendered routes as separate inputs rather than pretending this script covers them.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Or skip the browser setup
When your goal is to capture visual evidence of discovered pages rather than build the inventory itself, ScreenshotNeo provides a single-request website screenshot API. It can remove cookie/consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads and cache hits are not billed. Its MCP server gives Claude, Cursor and other MCP clients take_screenshot, get_page_info and capture_pdf tools.
For the full parameter list, see the ScreenshotNeo API documentation. Replace the target URL below with one from your inventory:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
You can also request full pages, selected elements, device presets, dark mode, custom CSS or JavaScript, waits, blocking rules, authentication headers, cookies, PDFs, signed links, asynchronous webhooks and bulk capture of up to 100 URLs per call. Every response includes X-Page-Verdict and X-Billed headers so you can see whether a clean shot was billed. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Rank #2
Troubleshooting common gaps
The sitemap contains fewer URLs than the crawl
Those extra URLs are crawl-only. Check whether they are intentionally excluded, parameter variants, pagination, media, or pages that should be added to the sitemap. Confirm that redirects and canonical targets are represented correctly.
Free tools Windows power users keep installed
One-click scans. No signup required.
The sitemap URL returns an error or HTML
Verify the fully qualified URL from robots.txt, follow sitemap-index nesting, check authentication and content type, and retain the failure in your audit log. A failed download is not evidence that the sitemap is empty.
Search Console shows a URL your crawler never reached
Inspect it. It may be an orphan, an old URL, a redirect, an authenticated route or a URL discovered through an external source. Compare the inspected canonical, crawl status and blocking details with your crawl record.
URLs differ only by parameters or case
Store the raw URL, then define a documented normalization policy for comparisons. Do not discard parameters until you know whether they change content, tracking, filtering or pagination.
JavaScript pages are missing
Use a rendering crawler, capture routes after interactions and compare the rendered link graph with the server HTML. Record that rendered discovery used a different method and date.
Recommended Free Tools
The crawl is slow or repeatedly fails
Lower concurrency, add backoff for transient failures, cache responses, set a maximum body size and split the work by host or path. Keep timeouts and error types in the export so “failed to fetch” is not confused with “URL does not exist.”
Performance, reliability and cost controls
- Start with cheap declarations: robots.txt and sitemaps reduce the initial crawl queue.
- Use a bounded crawl: set maximum URLs, depth, response size, concurrency and elapsed time.
- Preserve evidence: save response status, headers, canonical, noindex, source and timestamp for each record.
- Repeat incrementally: recrawl changed paths and compare normalized URL sets instead of rebuilding everything on every run.
- Separate environments: keep production, staging and authenticated hosts in distinct inventories.
- Budget rendering: browser rendering and screenshot capture cost more time and resources than fetching static HTML, so reserve them for JavaScript-dependent routes and visual checks.
A final URL count should always state what it counts: sitemap entries, unique discovered URLs, crawlable responses, Search Console-known URLs or indexed samples. Those numbers are not interchangeable.
FAQ
Can I export every URL Google knows in one file?
Not from the standard Page Indexing example list: Google documents a 1,000-URL limit for that interface view. Combine available Search Console data with sitemaps, crawling and targeted inspection instead.
Should an orphan URL always be removed?
No. First determine whether it is intentionally private, legacy, campaign-specific or required by users. Then choose a link, redirect, noindex or removal action that matches its purpose.
How should I report the result to another team?
Deliver the URL export with source flags, final responses, canonical and noindex fields, crawl date, host scope, authentication state and the rules used. That context makes the count reproducible.
Frequently Asked Questions
Can I export every URL Google knows in one file?
Not from the standard Page Indexing example list: Google documents a 1,000-URL limit for that interface view. Combine available Search Console data with sitemaps, crawling and targeted inspection instead.
Should an orphan URL always be removed?
No. First determine whether it is intentionally private, legacy, campaign-specific or required by users. Then choose a link, redirect, noindex or removal action that matches its purpose.
How should I report the result to another team?
Deliver the URL export with source flags, final responses, canonical and noindex fields, crawl date, host scope, authentication state and the rules used. That context makes the count reproducible.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The Bottom Line
For a defensible domain inventory, merge sitemap declarations, authenticated crawl results and Search Console evidence, then classify every URL by discovery, crawlability and indexing status. No single source can prove that it found every URL.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

