Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To audit every page, build a URL list from the XML sitemap, fetch each same-domain URL, extract the HTML <title> and <meta name="description"> values, render only pages whose metadata is injected by JavaScript, and export the raw values with quality flags. This workflow gives you coverage, provenance, and an editor-ready queue instead of a list of page titles with no explanation.

What you are extracting

The page title is the text inside the document’s <title> element. The meta description is the value of the content attribute on a tag such as <meta name="description" content="...">. Both belong in the document head. Keep the original strings as well as normalized copies: the original is needed for editing, while a normalized lowercase, whitespace-collapsed value makes duplicate detection reliable.

Plan the crawl before writing code

Define the URL boundary

  • Start with the site’s XML sitemap, usually /sitemap.xml, and follow any sitemap index files it references.
  • Normalize fragments and tracking parameters, then keep only permitted same-domain URLs.
  • Retain <lastmod> when supplied; it is useful for prioritizing recently changed pages.
  • If the sitemap is missing or incomplete, seed the crawl with the home page and follow canonical internal links. Record whether each URL came from a sitemap, an internal link, or a manual seed so omissions are explainable.

Decide what to record

For every request, store the requested URL, final URL after redirects, HTTP status, content type, fetch time, and whether the response was HTML. Also record the canonical URL and robots directives when available. Respect the site’s robots instructions and access controls; do not use a crawl as permission to bypass them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use bounded, repeatable fetching

Set a clear user agent, limit concurrency, retry transient failures with backoff, and deduplicate after normalization. These controls make a crawl dependable without assuming that every server has the same capacity.

Extract metadata from server-rendered HTML

When the response already contains the head markup, an HTTP client and an HTML parser are enough. This complete Python example reads a sitemap, follows sitemap indexes, fetches same-domain pages, and writes a CSV report.

import csv
import re
import time
from collections import defaultdict
from datetime import datetime, timezone
from urllib.parse import parse_qsl, urlencode, urljoin, urlparse, urlunparse
import requests
from bs4 import BeautifulSoup

START_SITEMAP = "https://example.com/sitemap.xml"
ALLOWED_HOST = urlparse(START_SITEMAP).netloc
USER_AGENT = "MetadataAudit/1.0 (+https://example.com/contact)"

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})

def normalize_url(raw):
    p = urlparse(raw.strip())
    query = [(k, v) for k, v in parse_qsl(p.query, keep_blank_values=True)
             if not (k.lower().startswith("utm_") or k.lower() in {"gclid", "fbclid"})]
    path = p.path or "/"
    return urlunparse((p.scheme.lower(), p.netloc.lower(), path, "", urlencode(query), ""))

def sitemap_urls(sitemap_url, seen=None):
    seen = seen or set()
    sitemap_url = normalize_url(sitemap_url)
    if sitemap_url in seen:
        return []
    seen.add(sitemap_url)
    response = session.get(sitemap_url, timeout=30)
    response.raise_for_status()
    soup = BeautifulSoup(response.content, "xml")
    if soup.find("sitemapindex"):
        found = []
        for loc in soup.select("sitemap > loc"):
            found.extend(sitemap_urls(loc.get_text(strip=True), seen))
        return found
    return [normalize_url(loc.get_text(strip=True))
            for loc in soup.select("url > loc")]

def clean(value):
    return re.sub(r"\s+", " ", value or "").strip()

def extract(html):
    soup = BeautifulSoup(html, "html.parser")
    title = clean(soup.title.get_text(" ", strip=True) if soup.title else "")
    tag = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "description"})
    description = clean(tag.get("content", "") if tag else "")
    canonical_tag = soup.find("link", rel=lambda v: v and "canonical" in v)
    canonical = canonical_tag.get("href", "").strip() if canonical_tag else ""
    robots_tag = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "robots"})
    robots = robots_tag.get("content", "").strip() if robots_tag else ""
    return title, description, canonical, robots

urls = []
for url in sitemap_urls(START_SITEMAP):
    if urlparse(url).netloc == ALLOWED_HOST:
        urls.append(url)
urls = sorted(set(urls))
rows = []
for url in urls:
    fetched_at = datetime.now(timezone.utc).isoformat()
    try:
        r = session.get(url, timeout=30, allow_redirects=True)
        content_type = r.headers.get("content-type", "")
        if "text/html" not in content_type.lower():
            title = description = canonical = robots = ""
            source = "not_html"
        else:
            title, description, canonical, robots = extract(r.text)
            source = "initial_html"
        rows.append({"url": url, "final_url": r.url, "status": r.status_code,
                     "content_type": content_type, "title_raw": title,
                     "description_raw": description, "metadata_source": source,
                     "canonical": canonical, "robots": robots,
                     "fetched_at": fetched_at})
    except requests.RequestException as exc:
        rows.append({"url": url, "final_url": "", "status": "error",
                     "content_type": "", "title_raw": "", "description_raw": "",
                     "metadata_source": "fetch_error", "canonical": "",
                     "robots": str(exc), "fetched_at": fetched_at})
    time.sleep(0.2)

for row in rows:
    row["title_normalized"] = clean(row["title_raw"]).lower()
    row["description_normalized"] = clean(row["description_raw"]).lower()
    flags = []
    if not row["title_raw"]: flags.append("missing_title")
    if not row["description_raw"]: flags.append("missing_description")
    if row["status"] != 200: flags.append("non_200")
    row["issue_flags"] = ";".join(flags)

groups = defaultdict(list)
for row in rows:
    if row["title_normalized"]: groups[("title", row["title_normalized"])].append(row)
    if row["description_normalized"]: groups[("description", row["description_normalized"])].append(row)
for row in rows:
    duplicate_groups = []
    for key, members in groups.items():
        if row in members and len(members) > 1:
            duplicate_groups.append(key[0] + ":" + key[1][:80])
    if duplicate_groups:
        row["duplicate_group"] = " | ".join(duplicate_groups)
        row["issue_flags"] += (";" if row["issue_flags"] else "") + "duplicate_metadata"
    else:
        row["duplicate_group"] = ""

fields = ["url", "final_url", "status", "content_type", "title_raw",
          "title_normalized", "description_raw", "description_normalized",
          "metadata_source", "canonical", "robots", "duplicate_group",
          "issue_flags", "fetched_at"]
with open("metadata-audit.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=fields)
    writer.writeheader()
    writer.writerows(rows)

Install the dependencies with python -m pip install requests beautifulsoup4. Replace the example domain and user-agent contact before running it. The script deliberately keeps failed requests in the output so a missing row cannot be mistaken for a page with no metadata.

Handle JavaScript-rendered metadata

A direct HTTP response can differ from the DOM a visitor sees. Some applications set or replace the title and description after JavaScript runs. Parse initial HTML first; send a URL to a browser-rendering queue only when a value is missing, when the site is known to inject metadata, or when a sample comparison shows a mismatch. Keep metadata_source as initial_html or rendered_dom.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright rendering example

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

def rendered_metadata(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle", timeout=60000)
        html = page.content()
        browser.close()
    soup = BeautifulSoup(html, "html.parser")
    title = soup.title.get_text(" ", strip=True) if soup.title else ""
    tag = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "description"})
    description = tag.get("content", "").strip() if tag else ""
    return title, description

Install the browser runtime with pip install playwright followed by playwright install chromium. Rendering every page is slower and more resource-intensive than parsing responses, so reserve it for the URLs that need it. Capture the final URL, status, and rendering timestamp just as you do for ordinary fetches.

Audit title and description quality

Missing and unusable values

  • Flag a missing title or description separately; fixing one does not fix the other.
  • Flag vague titles such as “Home” and titles that are excessively verbose for editorial review.
  • Check that the metadata describes the page’s visible main content rather than a template or unrelated section.

Duplicates and boilerplate

Group values after lowercasing and collapsing whitespace. Exact duplicates are obvious; near-duplicates need a review rule such as removing product IDs, dates, or punctuation before comparison. A site-wide suffix can be useful branding, but titles that differ only by that suffix or a token may still be indistinguishable. Descriptions should be written for their individual pages; identical or near-identical descriptions provide little help when different pages appear in search results.

Length without a false cutoff

Do not label a value “invalid” solely because it exceeds a universal character number. Search title links and snippets can be truncated to fit the display, so length is a review signal. Flag unusually short, unusually long, or empty values, then judge them against the page’s purpose and wording.

Robots and access findings

Keep blocked, excluded, redirected, and non-HTML responses in separate queues. A robots instruction can only be evaluated when the crawler can access the page, so do not infer a page’s directive from a failed request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Turn the export into an editorial queue

Use the CSV or a database row with these fields:

url, final_url, status, title_raw, title_normalized,
description_raw, description_normalized, metadata_source,
canonical, robots, lastmod, duplicate_group, issue_flags, fetched_at

Create separate worklists for missing metadata, duplicate or boilerplate values, JavaScript-only values, fetch errors, and metadata that does not match visible content. Include one sample URL for every duplicate group. Preserve raw values so an editor can see exactly what was served, not only the normalized comparison value.

Choose an implementation approach

Approach Best fit Strength Watch for
HTTP client plus parser Small site or one-off audit Simple, fast, inexpensive Misses metadata injected after load
Scrapy with an HTML parser Many pages or recursive discovery Crawling, extraction, and scheduling primitives You still need rendering logic for JavaScript pages
Browser renderer JavaScript-heavy applications Reads the post-load DOM Slower and more resource-intensive; use selectively
No-code crawler Teams needing scheduled reports Operational convenience Verify URL coverage, robots handling, and rendering on a sample

Compare tools on URL discovery, JavaScript rendering, duplicate detection, robots handling, retries and concurrency, export quality, scheduling, and total cost. The right choice depends on the site’s size and rendering behavior, not on a single feature checklist.

Troubleshoot common failures

The sitemap returns HTML or a 403

Check the response status and content type. The site may expose a sitemap index at another path, require authentication, or block your user agent. Confirm the permitted sitemap location with the site owner, then add an approved sitemap URL or a manual seed; do not silently treat an error page as an empty sitemap.

Many rows have blank titles

Inspect one raw response. If the head contains no title, the page may be genuinely missing metadata. If a browser shows a title, route that URL through the renderer and mark the source as rendered_dom. Also check that your parser is looking for the exact HTML element rather than a framework-specific data attribute.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Descriptions are missing but visible copy exists

Visible body text is not the same as a description meta tag. Flag the metadata as missing and write a page-specific description; do not automatically copy the first paragraph without editorial review.

The same URL appears several times

Normalize fragments, tracking parameters, hostname case, and trailing-slash policy before deduplication. Keep the final URL separately so redirects remain visible.

The crawl is slow or unreliable

Reduce concurrency, add bounded retries with exponential backoff, set connect and read timeouts, and log status, exception, and timestamp. Render only pages that failed the initial metadata test. Resume from the saved URL list rather than restarting discovery after every interruption.

Results disagree with search results

Your export shows the metadata served by the page at fetch time. Search systems may choose a different title link or snippet. Use the crawl to fix missing, duplicate, or misleading source metadata, then review search presentation separately.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo can capture pages through one request when you need a visual check alongside metadata work. It accepts the cookie or consent banner like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each cleanup step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed.

Use the API documentation at https://screenshotneo.com/docs/ for parameter details. The following call captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also provides an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools. Features include full-page and selector capture, device presets, retina scale, dark mode, custom CSS and JavaScript, clicks, waits, request blocking, cookies and headers, timezone and geolocation, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Should I crawl only canonical URLs?

Use sitemap and discovered URLs for coverage, but record the canonical value rather than discarding non-canonical pages. That lets you find duplicate or incorrectly canonicalized content.

Should a description be copied from the page body?

No. Body text can inform an editor, but the description should be a deliberate summary of that specific page.

Best Value
Sale
Latin Real Book: C Edition
  • Features Over 160 Latin Songs
  • Arranged for C Instruments
  • Standard Notation
  • 48 Pages

How often should the audit run?

Run it on a schedule that matches publishing activity, and run an additional crawl after a migration, template change, or metadata deployment.

Frequently Asked Questions

Should I crawl only canonical URLs?

Use sitemap and discovered URLs for coverage, but record the canonical value rather than discarding non-canonical pages. That lets you find duplicate or incorrectly canonicalized content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should a description be copied from the page body?

No. Body text can inform an editor, but the description should be a deliberate summary of that specific page.

How often should the audit run?

Run it on a schedule that matches publishing activity, and run an additional crawl after a migration, template change, or metadata deployment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.