What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable way to extract a website favicon is to inspect the page’s <link> elements, keep every icon candidate and its metadata, resolve each reference against the correct base URL, then inspect a linked web app manifest for additional icons. Do not assume that /favicon.ico exists or that one discovered image is the icon every browser, platform, or search result will use.

What you should extract

A useful favicon result is a candidate record, not just a URL. For every candidate, retain:

  • the source page URL;
  • whether the declaration came from HTML or a manifest;
  • the original href or src string;
  • the resolved absolute URL;
  • the relation tokens or manifest purpose;
  • declared type, sizes, and media values; and
  • the fetch result, such as HTTP status, content type, or an error.

This preserves enough provenance for a caller to choose an icon for a browser tab, app shortcut, search-oriented display, or another use case without pretending that a single universal selection algorithm exists.

How favicon discovery works

  1. Fetch the final page URL. Use the URL after redirects as the page base for resolving relative HTML references. Parse the document head’s link elements; do not probe only a conventional root path.
  2. Tokenize each rel value. Keep links whose relation tokens include icon. Also retain historical shortcut icon, apple-touch-icon, and apple-touch-icon-precomposed links because different consumers recognize different relations.
  3. Preserve selection hints. Save type, sizes, and media exactly as declared. These are hints used by browsers and platforms when several candidates are available; they are not a promise that every consumer will choose the same file.
  4. Resolve HTML references. Resolve a relative href against the page URL. An absolute URL may point to another host, including a CDN.
  5. Find a manifest. For <link rel="manifest" href="...">, resolve the manifest link against the page URL, fetch the JSON, and inspect its icons array.
  6. Resolve manifest icons against the manifest. A relative icon src is relative to the manifest file URL, not the HTML page URL. Keep src, sizes, type, and purpose when present.
  7. Fetch and validate separately. A declaration can be syntactically correct but inaccessible, blocked by policy, missing, or not an image. Record those outcomes instead of silently dropping the candidate.

Complete Python extractor (standard library)

The following script uses only Python’s standard library. It emits JSON records for HTML links and manifest icons, including fetch status and content type. It deliberately reports all candidates instead of selecting one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import sys
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urljoin
from urllib.request import Request, urlopen

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []
    def handle_starttag(self, tag, attrs):
        if tag.lower() == "link":
            self.links.append(dict(attrs))

def get_bytes(url, timeout=20):
    request = Request(url, headers={"User-Agent": "favicon-metadata-extractor/1.0"})
    try:
        with urlopen(request, timeout=timeout) as response:
            data = response.read()
            return {
                "ok": True,
                "status": getattr(response, "status", None),
                "content_type": response.headers.get_content_type(),
                "data": data,
                "error": None,
            }
    except (HTTPError, URLError, TimeoutError) as exc:
        return {"ok": False, "status": getattr(exc, "code", None), "content_type": None, "data": b"", "error": str(exc)}

def rel_tokens(value):
    return {token.lower() for token in (value or "").split()}

def public_result(source, original, resolved, attrs, fetch):
    return {
        "source": source,
        "original": original,
        "url": resolved,
        "rel": attrs.get("rel"),
        "type": attrs.get("type"),
        "sizes": attrs.get("sizes"),
        "media": attrs.get("media"),
        "purpose": attrs.get("purpose"),
        "fetch": {
            "ok": fetch["ok"],
            "status": fetch["status"],
            "content_type": fetch["content_type"],
            "error": fetch["error"],
        },
    }

def extract(page_url):
    page = get_bytes(page_url)
    if not page["ok"]:
        return {"page": page_url, "page_fetch": {k: page[k] for k in ("ok", "status", "content_type", "error")}, "candidates": []}
    html = page["data"].decode("utf-8", errors="replace")
    parser = LinkParser()
    parser.feed(html)
    candidates = []
    manifest_urls = []
    for attrs in parser.links:
        tokens = rel_tokens(attrs.get("rel"))
        href = attrs.get("href")
        if not href:
            continue
        if "icon" in tokens or "apple-touch-icon" in tokens or "apple-touch-icon-precomposed" in tokens:
            resolved = urljoin(page_url, href)
            candidates.append(public_result("html", href, resolved, attrs, get_bytes(resolved)))
        if "manifest" in tokens:
            manifest_urls.append(urljoin(page_url, href))
    for manifest_url in manifest_urls:
        manifest = get_bytes(manifest_url)
        if not manifest["ok"]:
            continue
        try:
            document = json.loads(manifest["data"].decode("utf-8", errors="replace"))
        except (UnicodeDecodeError, json.JSONDecodeError):
            continue
        for icon in document.get("icons", []):
            if not isinstance(icon, dict) or not icon.get("src"):
                continue
            src = icon["src"]
            resolved = urljoin(manifest_url, src)
            attrs = {"type": icon.get("type"), "sizes": icon.get("sizes"), "purpose": icon.get("purpose")}
            candidates.append(public_result("manifest", src, resolved, attrs, get_bytes(resolved)))
    return {"page": page_url, "page_fetch": {k: page[k] for k in ("ok", "status", "content_type", "error")}, "candidates": candidates}

if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python favicon_extract.py https://example.com/")
    print(json.dumps(extract(sys.argv[1]), indent=2))

Run it with python favicon_extract.py https://example.com/. The script treats a successful HTTP response and an image response as separate concerns: it records the server’s content type, but it does not claim that MIME metadata alone proves the bytes are a valid image. Add an image decoder if your application needs pixel-level validation.

Choosing among multiple candidates

Keep the full set until you know the consuming context. A practical ranking policy can consider the following fields, in this order:

Decision factor How to use it
Use context Tab, bookmark, installed web app, search display, and platform shortcut workflows may prefer different relations.
Fetch outcome Prefer candidates that were retrieved successfully; retain failures for diagnostics.
Declared dimensions Use sizes to avoid choosing a tiny asset when a larger suitable candidate exists. Treat any as a scalable or otherwise unspecified size rather than a numeric dimension.
Declared format Compare type with the response content type and the formats your consumer supports.
Media condition Evaluate media when your rendering environment has the information needed to do so.
Relation or purpose Use Apple touch relations or manifest purpose values for their intended platform or app context, not as a universal replacement for rel="icon".

There is no documented, cross-browser precedence algorithm that an independent scraper can reproduce exactly. Returning candidates plus metadata is safer than hard-coding one winner and presenting it as definitive.

Web app manifests: the second metadata source

A page can centralize application icons in a JSON manifest. The HTML link to that manifest is resolved against the page URL. Each manifest icon’s src is then resolved against the manifest URL, which matters when the manifest lives in a subdirectory or on another origin.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Manifest processing has security and deployment constraints. Cross-origin manifests can require appropriate CORS handling under the manifest processing rules, and icon retrieval is subject to the manifest owner’s img-src Content Security Policy. Consequently, “declared in markup” and “retrievable by your service” are different states. Record both.

Manifest entries may include purpose, such as a normal icon or a maskable icon. Preserve the value for the installer or renderer that understands it; do not infer that every browser will honor it identically.

Platform and search-result caveats

Google’s favicon guidance describes eligibility for Google Search, not a universal browser algorithm. It recognizes standard icon declarations and historical shortcut relations, and its documentation warns that display is not guaranteed. After a site changes an icon, Google says recrawling and processing can take several days to several weeks, depending on how its systems decide that a refresh is needed.

Apple platforms use apple-touch-icon for certain web-clip or startup-placeholder uses rather than the ordinary rel="icon" pathway. That is why an extractor should retain the exact relation token instead of normalizing every icon to one generic category.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not substitute a brand logo, Open Graph image, or social preview for a favicon unless your product explicitly labels that fallback. Those metadata types are not established as equivalent to icon relations or manifest icons.

Common failures and fixes

No candidates found

Cause: the page has no qualifying link, the response was not the final HTML document, or parsing stopped before the head was received. Fix: follow redirects, save the final response for inspection, and check for a manifest link. Do not conclude that the site has no icon merely because /favicon.ico returned an error.

Relative URL points to the wrong file

Cause: resolving a manifest src against the page URL. Fix: resolve the manifest link against the page first, then resolve every icon src against that manifest URL.

HTTP 403, 401, or a bot-check page

Cause: the host requires credentials, blocks automated clients, or returned an interstitial instead of the icon. Fix: record the status, avoid treating the response as an image, and use an authorized request only when you have permission. A user-agent change is not a bypass for access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

HTML declares an icon but fetching fails

Cause: cross-origin policy, network restrictions, a restrictive content security policy, DNS failure, timeout, or a deleted asset. Fix: keep the declaration in your output, store the fetch error, and retry with bounded timeouts and backoff where appropriate.

The downloaded bytes are not an image

Cause: a login page, error document, or HTML bot challenge was served at the icon URL. Fix: inspect status and content type, then validate the file with an image parser before handing it to downstream code.

Several icons appear equally suitable

Cause: the site intentionally publishes multiple dimensions, formats, media conditions, or platform-specific relations. Fix: return all candidates and let the caller select for its target context; document your selection rule if you must choose one.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost controls

  • Fetch the page once and reuse its parsed link list.
  • Fetch a manifest only when one is linked, and deduplicate identical resolved URLs.
  • Set finite connect and read timeouts; one unresponsive icon should not block the entire page result.
  • Cap response sizes before storing bytes, especially for URLs that unexpectedly return large documents.
  • Cache successful metadata with a revalidation policy appropriate to your application, but keep the source URL and retrieval time.
  • Apply concurrency limits when processing many pages so your client does not overload target sites.
  • Keep status, content type, redirect destination, and error details for reproducible debugging.
  • Respect robots, authentication, rate limits, and terms that apply to the sites you fetch.

Discovery itself is inexpensive, but a robust service pays for DNS, TLS, transfer, parsing, image validation, retries, and storage. Measuring those stages separately helps you set useful quotas and prevents repeated downloads of the same icon.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If you also need a visual capture of the page while collecting metadata, ScreenshotNeo is the first screenshot API to try: it removes common consent banners, popups, and chat widgets before capture, bills only clean shots, and its lowest paid plan starts at $5.

It is not a replacement for parsing favicon declarations; use the extraction workflow above when you need URLs and metadata. Use ScreenshotNeo when a rendered page image or PDF is part of the job, without maintaining a headless-browser stack.

One GET request returns the rendered asset. See the ScreenshotNeo API documentation for the complete parameter set.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o page.webp

Equivalent Python and Node.js calls are:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("page.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('page.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo’s response identifies the page verdict and billing outcome in X-Page-Verdict and X-Billed headers. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.