Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To find links in an HTML string with PHP, parse the markup with PHP’s DOM extension, select every <a> element, and read its href attribute. This preserves the document structure and avoids treating HTML as a regular expression problem.

<?php
$html = '<a href="https://example.com">Example</a>';

$dom = new DOMDocument();
$dom->loadHTML($html);

$links = [];
foreach ($dom->getElementsByTagName('a') as $anchor) {
    $links[] = $anchor->getAttribute('href');
}

print_r($links);

The result is an array of href values, in document order. The examples below show how to handle strings and files, choose the parser for your PHP version, filter or deduplicate results, resolve relative URLs, preserve useful metadata, and diagnose empty or unexpected output.

What “all links” means in this example

This method extracts URLs represented by href attributes on anchor elements. It does not search JavaScript, visible text, CSS, images, arbitrary attributes, or every URL-shaped string in a page. If your application also needs other HTML links, query their elements explicitly: for example, <area href>, <link href>, or <iframe src>.

The parser does not fetch URLs, verify that they work, canonicalize them, remove duplicates, or automatically convert relative references into absolute URLs. Those are separate choices your application must make.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract href values from an HTML string

Compatible DOMDocument version

DOMDocument::loadHTML() is the broadly compatible approach for existing projects and older PHP runtimes. It returns a DOM tree; getElementsByTagName('a') returns a DOMNodeList that can be iterated directly.

<?php
declare(strict_types=1);

$html = <<<'HTML'
<!doctype html>
<html>
  <body>
    <a href="/docs">Documentation</a>
    <a href="https://example.com/pricing">Pricing</a>
    <a href="mailto:support@example.com">Email support</a>
    <a href="">Current page</a>
    <a>No href attribute</a>
  </body>
</html>
HTML;

$previous = libxml_use_internal_errors(true);
$dom = new DOMDocument();
$dom->loadHTML($html, LIBXML_NONET);
libxml_clear_errors();
libxml_use_internal_errors($previous);

$links = [];
foreach ($dom->getElementsByTagName('a') as $anchor) {
    if (!$anchor->hasAttribute('href')) {
        continue;
    }
    $links[] = $anchor->getAttribute('href');
}

var_export($links);

Checking hasAttribute() lets you distinguish a missing href from an explicitly empty one. Remove that check if your requirements are to keep both cases as empty strings. The LIBXML_NONET flag prevents network access by libxml while parsing this in-memory string.

HTML5 parsing on PHP 8.4 and later

PHP 8.4 introduced DomHTMLDocument, intended for HTML5-conforming parsing. PHP’s documentation recommends it for modern HTML instead of relying on DOMDocument::loadHTML(), whose tree construction follows an HTML 4 parser and can differ from a browser on malformed markup.

<?php
declare(strict_types=1);

$html = '<main><a href="/news">News</a></main>';
$document = DomHTMLDocument::createFromString($html);

$links = [];
foreach ($document->getElementsByTagName('a') as $anchor) {
    if ($anchor->hasAttribute('href')) {
        $links[] = $anchor->getAttribute('href');
    }
}

var_export($links);

Use this version only when your minimum runtime is PHP 8.4 or newer and the DOM extension provides the class. Keep the legacy example for applications that must support earlier versions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read links from an HTML file

For a local file, load the file into the DOM and use the same iteration. Check the return value so a missing or unreadable path is reported instead of producing a misleading empty list.

<?php
declare(strict_types=1);

$path = __DIR__ . '/page.html';
if (!is_readable($path)) {
    throw new RuntimeException("Cannot read {$path}");
}

$previous = libxml_use_internal_errors(true);
$dom = new DOMDocument();
$loaded = $dom->loadHTMLFile($path, LIBXML_NONET);
$errors = libxml_get_errors();
libxml_clear_errors();
libxml_use_internal_errors($previous);

if ($loaded === false) {
    throw new RuntimeException('The file could not be parsed as HTML.');
}

foreach ($dom->getElementsByTagName('a') as $anchor) {
    if ($anchor->hasAttribute('href')) {
        echo $anchor->getAttribute('href'), PHP_EOL;
    }
}

loadHTMLFile() reads the path itself; if the HTML is already in a variable, use loadHTML() instead. Parsing warnings are not necessarily fatal, because real-world HTML is often incomplete, so decide whether your application should log them, reject the document, or continue.

Keep useful link data instead of only the URL

Often you need the anchor text, target, or other attributes for reporting. Store a record for each anchor while traversing the same node list.

<?php
$records = [];
foreach ($dom->getElementsByTagName('a') as $anchor) {
    if (!$anchor->hasAttribute('href')) {
        continue;
    }

    $records[] = [
        'href'   => trim($anchor->getAttribute('href')),
        'text'   => trim($anchor->textContent),
        'target' => $anchor->getAttribute('target'),
        'rel'    => $anchor->getAttribute('rel'),
    ];
}

print_r($records);

Trimming is an application policy, not an automatic DOM behavior. Preserve the original value when whitespace is meaningful to your workflow, or keep both the raw and normalized values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Filter, deduplicate, and classify results

Ignore empty and fragment-only references

<?php
$links = [];
foreach ($dom->getElementsByTagName('a') as $anchor) {
    if (!$anchor->hasAttribute('href')) {
        continue;
    }

    $href = trim($anchor->getAttribute('href'));
    if ($href === '' || str_starts_with($href, '#')) {
        continue;
    }
    $links[] = $href;
}

Remove duplicates without changing order

<?php
$uniqueLinks = array_values(array_unique($links, SORT_STRING));

This compares the strings exactly. Thus, a relative path and its absolute equivalent, or URLs that differ only by a fragment, remain separate until you normalize them according to your own rules.

Keep only HTTP(S) links

<?php
$httpLinks = array_values(array_filter(
    $links,
    static function (string $href): bool {
        $scheme = strtolower((string) parse_url($href, PHP_URL_SCHEME));
        return in_array($scheme, ['http', 'https'], true);
    }
));

This deliberately excludes mail, telephone, JavaScript, data, fragment, and relative references. Validate any URL further before using it in a request or redirect.

Resolve relative href values against a base URL

HTML commonly contains /about, ../contact, or pricing rather than a complete URL. PHP’s DOM parser returns those strings as written; it does not know which page URL should be their base. Resolve them only when you have the correct document URL.

<?php
function absoluteUrl(string $href, string $base): ?string
{
    $href = trim($href);
    if ($href === '') {
        return null;
    }

    if (parse_url($href, PHP_URL_SCHEME) !== null) {
        return $href;
    }

    $baseParts = parse_url($base);
    if (!$baseParts || empty($baseParts['scheme']) || empty($baseParts['host'])) {
        throw new InvalidArgumentException('Base URL must include a scheme and host.');
    }

    $origin = $baseParts['scheme'] . '://' . $baseParts['host'];
    if (isset($baseParts['port'])) {
        $origin .= ':' . $baseParts['port'];
    }

    if (str_starts_with($href, '//')) {
        return $baseParts['scheme'] . ':' . $href;
    }
    if (str_starts_with($href, '/')) {
        return $origin . $href;
    }

    $basePath = $baseParts['path'] ?? '/';
    $directory = rtrim(str_replace('\', '/', dirname($basePath)), '/');
    $path = ($directory === '' ? '' : $directory) . '/' . $href;
    $segments = [];
    foreach (explode('/', $path) as $segment) {
        if ($segment === '' || $segment === '.') {
            continue;
        }
        if ($segment === '..') {
            array_pop($segments);
        } else {
            $segments[] = $segment;
        }
    }

    return $origin . '/' . implode('/', $segments);
}

$base = 'https://example.com/docs/index.html';
$absolute = [];
foreach ($links as $href) {
    $url = absoluteUrl($href, $base);
    if ($url !== null) {
        $absolute[] = $url;
    }
}

This compact helper covers common paths and protocol-relative references. URL resolution has additional edge cases, including query-only references, fragments, and encoded path segments; use a standards-compliant URI resolver when those cases affect correctness.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing and encoding pitfalls

HTML5 versus the legacy parser

DOMDocument::loadHTML() uses HTML 4 parsing rules. Browsers use HTML5 rules, so malformed or modern markup can produce a different tree. On PHP 8.4 and later, prefer DomHTMLDocument for modern documents. Do not describe either parser as an HTML sanitizer: PHP specifically warns that loadHTML() cannot safely sanitize untrusted HTML.

Character encoding

The DOM extension works with UTF-8. If the source is in another encoding, convert it based on the source’s actual declaration before parsing. PHP identifies mb_convert_encoding(), UConverter::transcode(), and iconv() as possible conversion tools.

<?php
// Example only: use the encoding you have verified for the source.
$utf8Html = mb_convert_encoding($html, 'UTF-8', 'ISO-8859-1');
$dom = new DOMDocument();
$dom->loadHTML($utf8Html, LIBXML_NONET);

Guessing an encoding can corrupt link text or attribute values. Preserve the original bytes when you need forensic accuracy and record the conversion separately.

Common failures and fixes

The result is an empty array

  • Confirm the input actually contains <a href="..."> elements. Links inserted later by JavaScript are not present in a static HTML string.
  • Check that the DOM extension is enabled in the PHP runtime running your script.
  • Verify the file path and permissions when using loadHTMLFile().
  • Inspect parser warnings with libxml_get_errors() and ensure the input encoding is handled.

Only some links appear

Inspect whether the missing items are in scripts, comments, SVG, or non-anchor elements. Query each required element and attribute explicitly; an anchor-only query is not a site-wide URL crawler.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Characters look corrupted

Convert non-UTF-8 input before parsing, using the source’s declared or otherwise verified encoding. Do not fix this by blindly applying a single conversion to every document.

Warnings are printed before my output

Use libxml_use_internal_errors(true) around parsing, collect and log the errors, then restore the previous libxml setting. Suppressing output without examining the errors can hide malformed input or a wrong file path.

Untrusted markup is being “cleaned”

Extraction and sanitization are different jobs. The DOM parser can build a tree for inspection, but DOMDocument::loadHTML() is not a safe sanitizer. Use a sanitizer designed for your threat model before rendering untrusted content.

Performance, memory, and reliability choices

  • A DOM parser builds a tree in memory, which makes traversal straightforward but means very large documents consume corresponding memory. Enforce input-size limits before parsing untrusted or remote content.
  • For a local batch, parse each file once and process its nodes during that pass instead of repeatedly reparsing the same string.
  • Use LIBXML_NONET when parsing supplied markup so external references are not fetched by libxml.
  • Separate extraction from network checks. First collect and normalize links; then apply request timeouts, redirect limits, rate limits, and robots or access policies appropriate to your crawler.
  • Keep duplicates when occurrence counts or page-position analysis matters; deduplicate only for a set of destinations.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your real goal is to obtain a rendered page image or PDF after discovering a URL, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL in one request and can return PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a direct request, see the ScreenshotNeo API documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same endpoint can be called from PHP, Python, or Node.js:

<?php
$url = 'https://api.screenshotneo.com/v1/shot';
$query = http_build_query([
    'access_key' => 'YOUR_API_KEY',
    'url' => 'https://stripe.com',
]);
$data = file_get_contents($url . '?' . $query);
file_put_contents('shot.webp', $data);
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also supports full-page and element captures, device and viewport settings, dark mode, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, PDF controls, caching, signed links, asynchronous jobs, bulk capture, usage reporting, and an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI clients. Every feature is included on every plan: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000, with yearly billing providing two months free. Create a free ScreenshotNeo account.

Choosing an implementation

Requirement Recommended approach Reason
Support older PHP versions DOMDocument::loadHTML() Longstanding API, with HTML 4 parsing limitations documented.
HTML5-conforming parsing DomHTMLDocument on PHP 8.4+ Designed for modern HTML parsing behavior.
Extract local markup Parse once, iterate a nodes Simple, deterministic, and does not perform network requests.
Obtain links from a rendered, JavaScript-heavy page Render the page first, then parse the resulting HTML Static source may not contain links created in the browser.

Frequently Asked Questions

Can I use a regular expression to find every HTML link?

Use a DOM parser for HTML structure. A regular expression can miss malformed markup, quoted-attribute variations, nested elements, and entities.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does this code crawl every page linked from the document?

No. It extracts attribute values from one HTML document. Following links requires a separate crawler with URL, access, rate, and termination policies.

How can I include links inside an

or element?

Run additional DOM queries for those tag names and read their relevant attributes, such as href; do not assume the anchor query includes them.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.