The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →To generate an XML sitemap, export the canonical, indexable URLs on your site into a UTF-8 XML file, publish it at a stable address such as /sitemap.xml, validate it, and submit it in Google Search Console. Small sites can use a text editor; CMSs usually generate one automatically; larger sites should generate it from the database or application routes so it stays current. A sitemap helps search engines discover URLs, but submission does not guarantee crawling or indexing.
What an XML sitemap does
An XML sitemap is a machine-readable list of URLs, with optional metadata about updates. It can also describe video, image, and news content and the relationships among sitemap files. Its primary job is discovery: it gives search engines another way to find pages that may be difficult to reach through internal links.
Google’s guidance says a sitemap is especially useful for large sites, new sites with few external links, and sites whose important content is video, image, or news material. A site with about 500 pages or fewer that is comprehensively linked and has little specialized media may not need one. In every case, a sitemap is a hint rather than a ranking or indexing guarantee.
Choose a generation method
CMS-generated sitemap
WordPress, Wix, Blogger, and comparable platforms commonly create a sitemap or sitemap index automatically. Find the platform’s documented sitemap URL and settings before installing another generator. Check whether the output includes only canonical, indexable URLs and whether taxonomy, author, search, or attachment pages should be excluded.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Manual sitemap for a small site
For fewer than a few dozen stable URLs, a text editor is practical. You must update the file whenever canonical URLs change, pages are removed, or important content is added. Manual maintenance becomes error-prone as soon as several people or systems publish pages.
Application or database generation
For a larger or frequently changing site, generate the sitemap from the same database or routing layer that defines canonical pages. This gives the generator authoritative URLs, makes scheduled updates possible, and lets you filter redirects, duplicates, and non-indexable records before writing XML. A crawler can discover URLs, but a database or CMS export usually gives better control over canonical status and update dates.
What to compare before choosing
- URL authority: database or CMS data is normally more reliable than a crawl of navigation.
- Automation: determine whether generation runs on publish, on a schedule, or only when you edit a file.
- Filtering: verify canonical, redirect, noindex, duplicate, staging, and parameterized URLs are handled deliberately.
- Scale: confirm support for sitemap indexes and the protocol limits below.
- Extensions: check whether image, video, or news metadata is required for your content.
- Operations: look for validation, deployment, logging, and Search Console monitoring.
- Last modification control: use a source that can provide a trustworthy significant-update timestamp.
XML sitemap structure and hard limits
Write UTF-8 XML with fully qualified absolute URLs. A minimal file looks like this:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
</url>
<url>
<loc>https://example.com/docs/getting-started</loc>
<lastmod>2026-09-20</lastmod>
</url>
</urlset>
Each <loc> must be an absolute URL on the intended site. Escape XML-sensitive characters in every value: for example, write an ampersand as &. Include only URLs you want considered for search, normally the canonical HTTPS version that returns a successful response.
Google’s current limits are 50,000 URLs or 50 MB uncompressed per sitemap. When either limit is reached, split the files and publish a sitemap index:
Rank #2
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemaps/products-1.xml</loc>
</sitemap>
<sitemap>
<loc>https://example.com/sitemaps/products-2.xml</loc>
</sitemap>
</sitemapindex>
Keep every child file within both limits. URL order does not matter to Google. Compression can reduce transfer size, but the 50 MB limit is measured before compression.
Use optional tags carefully
<lastmod> is useful only when it is consistently accurate and reflects a significant page update. Do not change it merely because a copyright year changed or a deployment ran. Google ignores <priority> and <changefreq>, so adding them creates maintenance without improving processing.
How to build a sitemap scraper or generator
A “sitemap scraper” can mean a crawler that collects links or a generator that exports known URLs. For search submission, make the generator’s final inventory canonical and indexable rather than blindly copying every link a crawler sees.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Choose the source. Query the CMS/database or crawl the public site when no authoritative export exists.
- Normalize URLs. Resolve relative links, use the preferred scheme and host, remove tracking parameters, and normalize trailing-slash rules consistently.
- Filter records. Exclude redirects, errors, duplicate URLs, noindex pages, blocked staging content, login-only pages, and non-canonical variants unless you have a deliberate reason to include one.
- Check reachability. Request each URL and record the HTTP response. A sitemap entry should normally resolve successfully and represent the canonical page.
- Deduplicate and sort. Stable ordering makes diffs and troubleshooting easier, although Google does not require a particular order.
- Write escaped UTF-8 XML. Generate one or more sitemap files and an index when needed.
- Validate before deployment. Parse the XML, check namespace and limits, and inspect a sample of URLs and timestamps.
- Publish atomically. Write a temporary file, validate it, then replace the public file so readers never see half-written XML.
- Schedule regeneration. Run on publication events or a dependable schedule, and log counts, exclusions, failures, and the generation time.
Minimal generation pseudocode
records = load_canonical_indexable_pages()
urls = []
for record in records:
url = normalize(record.canonical_url)
if record.noindex or record.redirect or not url:
continue
if url not in urls:
urls.append(url)
write_sitemaps(urls, max_urls=50000, max_uncompressed_bytes=50_000_000)
In production, measure the serialized byte size while writing, not just the number of URLs. XML escaping can make the output larger than the source strings. If you support image, video, or news extensions, validate their namespaces and required fields separately.
Publish, test, and submit the file
- Place the file at a stable URL, preferably at the site root, such as
https://example.com/sitemap.xml. A root location makes it easier for the file to cover the site’s paths. - Fetch the public URL over HTTPS and confirm the response is successful, the content is XML, and no login, HTML error page, or redirect chain is returned.
- Open the file and verify that every
<loc>is absolute, belongs to the intended site, and matches a canonical page. - Run an XML parser and a sitemap validator. Also inspect HTTP responses for a representative sample and for every failed URL reported by your generator.
- In Google Search Console, open the Sitemaps report, enter the sitemap or sitemap-index URL, and submit it. You can also advertise it in
robots.txtwith a line such asSitemap: https://example.com/sitemap.xml. - For automated operations, the Search Console API can submit sitemaps programmatically. Continue monitoring the Sitemaps report for fetch and processing errors.
Submission is not a command to crawl. Google’s documented caveat is that submitting a sitemap merely provides a hint; Google may not download it or use it for crawling the listed URLs.
Rank #3
- New
- Mint Condition
- Dispatch same day for order received before 12 noon
- Guaranteed packaging
- No quibbles returns
Common failures and fixes
“Could not fetch” in Search Console
Fetch the exact public URL from outside your application network. Check DNS, TLS, authentication, firewall rules, robots handling, server timeouts, and whether the URL returns an HTML error page or an unexpected redirect. Fix the public response, then resubmit.
Invalid XML or parse errors
Look for unescaped ampersands, unclosed tags, a missing XML namespace, invalid control characters, or a file saved in a non-UTF-8 encoding. Parse the deployed file, not only the local source, because a proxy or build step may alter it.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →URLs are excluded or not indexed
Compare each URL with its canonical tag, robots directives, HTTP status, and rendered content. Remove redirects, duplicates, noindex pages, and URLs blocked by an intentional policy. A valid sitemap cannot override those signals and cannot guarantee indexing.
Too many URLs or an oversized file
Split at or below 50,000 URLs and 50 MB uncompressed, then create a sitemap index containing absolute links to every child file. Track both counters during generation.
Stale or misleading lastmod dates
Source dates from the record’s actual significant content update. If your system cannot provide reliable values, omit <lastmod> rather than refreshing it on every build.
New pages are missing
Check the generator’s source query, publication filters, cache, and scheduled job. Compare the generated URL count with the CMS or database count, and alert when the count drops unexpectedly.
Performance, reliability, and cost decisions
Database generation is usually faster and less resource-intensive than crawling your own site, while crawling can find orphaned public URLs that your content database does not know about. For very large inventories, stream output, shard work by content type, and avoid loading every URL into memory. Cache stable records, but invalidate the cache when canonical or indexability fields change.
Use retries with limits for transient HTTP checks, record failures rather than silently dropping them, and publish the last known-good file if a new build fails validation. Keep generation logs containing timestamp, URL count, byte size, excluded-count breakdown, and error samples. These controls cost less operationally than discovering a broken sitemap after a release.
There is no requirement to buy a separate sitemap service. A CMS feature, scheduled application job, or small script is sufficient; paid crawler or validator software is most useful when you need broader technical auditing, not merely XML output.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not an XML sitemap generator. It can nevertheless help you visually check representative pages after your sitemap job publishes them, without maintaining browser automation. One GET request returns a PNG, JPEG, WebP, or PDF; cookie and consent banners, newsletter popups, and chat widgets are removed before capture.
Use the API documented at https://screenshotneo.com/docs/:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed as clean shots, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Does every website need a sitemap?
No. A small, well-linked site with little specialized media may gain little from one, although maintaining a valid sitemap is harmless when it is automated.
Can I submit multiple sitemap files?
Yes. Submit an index that lists the child files, or submit individual files, while keeping each file within the protocol limits.
Should URLs be sorted by importance?
No. Google does not require a URL order, so sort for operational clarity rather than ranking.
Will a sitemap make a page rank or get indexed?
No. It supports discovery only; crawling and indexing depend on Google’s other systems and your page’s technical and content signals.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

