Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use PHP’s native DOMDocument and DOMXPath to parse HTML and select class names safely. The XPath expression below treats class as a whitespace-separated token list, so it matches class="card featured" without accidentally matching class="cardinal".

Find every element with a class in native PHP

This complete example parses an HTML string, creates an XPath evaluator, selects every element containing the card class, and prints its text:

<?php
$html = '<div class="card featured">A</div><div class="card">B</div>';

$dom = new DOMDocument();
libxml_use_internal_errors(true);
$dom->loadHTML($html);
$xpath = new DOMXPath($dom);

$nodes = $xpath->query(
    "//*[contains(concat(' ', normalize-space(@class), ' '), ' card ')]"
);

foreach ($nodes as $node) {
    echo trim($node->textContent), PHP_EOL;
}

The output is:

A
B

DOMDocument builds a document tree, while DOMXPath evaluates XPath 1.0 expressions against that tree. The normalize-space() and concat() calls add spaces around the normalized class attribute. That makes card an exact token rather than a substring.

Why an exact class-token query matters

HTML permits multiple classes in one attribute and permits arbitrary whitespace between them. A test such as //*[@class='card'] only matches an attribute whose entire value is exactly card; it misses class="card featured". A test such as contains(@class, 'card') has the opposite problem: it also matches names such as cardinal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use this reusable predicate whenever you need one class token:

contains(concat(' ', normalize-space(@class), ' '), ' CLASS_NAME ')

Replace CLASS_NAME with the token you want, including the surrounding spaces shown in the expression. If the source has no class attribute, the predicate simply evaluates to false.

Useful DOMXPath selectors

Restrict the class to a tag

To find only links with the button class, change the wildcard element test to a:

$nodes = $xpath->query(
    "//a[contains(concat(' ', normalize-space(@class), ' '), ' button ')]"
);

Require a class inside a particular ancestor

This expression finds price elements below an element with the product class:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
$nodes = $xpath->query(
    "//*[contains(concat(' ', normalize-space(@class), ' '), ' product ')]" .
    "//*[contains(concat(' ', normalize-space(@class), ' '), ' price ')]"
);

For a simpler descendant relationship, select the descendant directly from a known tag or class. Keep the token-safe predicate on each class rather than switching to a substring test.

Combine several conditions

Use XPath’s and operator when an element must have two class tokens:

$nodes = $xpath->query(
    "//*[contains(concat(' ', normalize-space(@class), ' '), ' card ')" .
    " and contains(concat(' ', normalize-space(@class), ' '), ' featured ')]"
);

Select one expected element safely

query() returns a DOMNodeList, even when you expect one result. Check its length before reading item zero:

$nodes = $xpath->query(
    "//*[@id='checkout']"
);

if ($nodes->length === 0) {
    echo "Checkout element not found", PHP_EOL;
} else {
    $checkout = $nodes->item(0);
    echo trim($checkout->textContent), PHP_EOL;
}

This avoids treating a missing match as a valid node. If more than one match is possible, iterate the whole list instead of silently discarding the rest.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read text, attributes, and HTML from each match

Text content

Use textContent for the text inside an element, including text nested in child elements. Trimming is useful when source indentation contributes whitespace:

foreach ($nodes as $node) {
    $text = trim($node->textContent);
    echo $text, PHP_EOL;
}

Attributes

Check that an attribute exists before using it:

foreach ($nodes as $node) {
    $href = $node instanceof DOMElement
        ? $node->getAttribute('href')
        : '';

    echo $href, PHP_EOL;
}

getAttribute() returns an empty string when the attribute is absent, so apply any required validation before treating the value as a URL, identifier, or other required field.

Keep the node for later DOM operations

Each result remains a DOM node. You can inspect nodeName, walk child nodes, or use another XPath query with that node as the context. The initial query does not modify the document.

Parsing real input reliably

Suppress and inspect parser warnings

HTML from the web is often incomplete or malformed. loadHTML() attempts to repair it and can emit libxml warnings. The example enables internal error handling so warnings do not leak into normal output. In a production script, clear the error buffer after parsing and log the errors you actually need to diagnose.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
libxml_use_internal_errors(true);
$dom = new DOMDocument();
$ok = $dom->loadHTML($html, LIBXML_NOERROR | LIBXML_NOWARNING);
$errors = libxml_get_errors();
libxml_clear_errors();

if ($ok === false) {
    throw new RuntimeException('The HTML could not be parsed.');
}

Encoding

DOMDocument::loadHTML() parses the string you provide; it does not solve every encoding problem. If text is garbled, verify the response bytes and the document’s declared encoding before parsing. Do not assume that fetching a URL, authenticating, or decoding compressed content is part of DOM parsing.

Remote pages and JavaScript

Fetching a remote page is a separate operation. Supply the downloaded HTML to loadHTML() only after handling redirects, authentication, timeouts, and response encoding in your HTTP client. The DOM extension parses the HTML present in that string; it does not run browser JavaScript. Elements inserted after page load by client-side code will not be available unless you use a browser renderer or another service that executes the page.

Use Symfony DomCrawler for concise CSS selectors

When Composer is available, Symfony DomCrawler provides a shorter, chainable API for navigating HTML and XML documents. Install the crawler and its CSS-selector dependency:

composer require symfony/dom-crawler symfony/css-selector

Then select the class with the familiar CSS syntax:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
require __DIR__ . '/vendor/autoload.php';

use SymfonyComponentDomCrawlerCrawler;

$html = '<div class="card featured">A</div><div class="card">B</div>';
$crawler = new Crawler($html);

foreach ($crawler->filter('.card') as $element) {
    echo trim($element->textContent), PHP_EOL;
}

filter('.card') returns a new Crawler containing all matching elements. You can chain selectors and use helper methods such as text(), attr(), extract(), and each():

$prices = $crawler->filter('.product .price')->each(
    fn (Crawler $node) => $node->text('')
);

$links = $crawler->filter('a.button')->each(
    fn (Crawler $node) => [
        'label' => trim($node->text('')),
        'href' => $node->attr('href'),
    ]
);

Pass a default to text() when no match is acceptable. Without a default, calling text() on an empty selection throws an exception. For structural conditions that CSS cannot express conveniently, use filterXPath(); Symfony supports both selector styles.

Choose DOMXPath or DomCrawler

Requirement DOMDocument + DOMXPath Symfony DomCrawler
Dependencies Native PHP DOM APIs; no Composer package Composer packages: symfony/dom-crawler and symfony/css-selector
Ordinary class selection Longer token-safe XPath predicate Concise CSS such as .card
Structural and attribute predicates XPath is expressive and direct Use CSS where convenient or filterXPath()
Result handling DOMNodeList; iterate and read DOM properties New Crawler instances; chain and use extraction helpers

Choose the native approach for a small script or a project that must avoid third-party dependencies. Choose DomCrawler when readable CSS selectors, chaining, and extraction helpers make the Composer dependency worthwhile. Neither approach is a browser; both operate on the markup available to the parser.

Troubleshoot missing or unexpected matches

Symptom Likely cause Fix
No class matches The class is added by JavaScript after the original HTML loads. Capture or obtain the rendered DOM, or use a browser-capable workflow before parsing.
card misses card featured The XPath used an exact attribute comparison. Use the contains(concat(' ', normalize-space(@class), ' '), ' card ') token test.
Too many matches A substring test such as contains(@class, 'card') also matches longer names. Use the token-safe predicate and add tag, ancestor, or additional-class conditions.
Warning output corrupts a JSON or HTML response libxml warnings are being printed during parsing. Enable internal errors, suppress or log warnings, and clear the libxml error buffer.
“Class not found” at runtime The PHP DOM extension is unavailable in the installed build. Enable the DOM/XML extension for that PHP installation, then verify with class_exists('DOMDocument').
DomCrawler selector fails One or both Symfony packages are missing from the autoloader. Run the Composer install command, include vendor/autoload.php, and confirm the packages are present.
Unexpected whitespace in text HTML indentation and nested nodes are part of textContent. Trim the value and normalize whitespace according to your application’s data rules.

Performance, reliability, and security notes

  • Parse once and reuse the same DOMXPath or Crawler for several selectors instead of reparsing identical HTML.
  • Use a narrower XPath (for example, a specific tag or ancestor) when the document is large and you do not need every element.
  • For a single expected result, check the collection before reading its first item; this makes changed or incomplete markup an explicit case.
  • Treat downloaded HTML as untrusted input. Validate URLs, attributes, and extracted text before storing, displaying, or using them in another request.
  • Record the source response and parser errors when a selector unexpectedly returns zero results. A selector cannot find markup that was never present in the input string.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is to obtain a clean screenshot of a page before inspecting its HTML visually, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request returns PNG, JPEG, WebP, or PDF. The full option set includes full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work.

Use the API key and target URL shown below; the ScreenshotNeo documentation covers the available options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Claude, Cursor, and other MCP clients can use ScreenshotNeo’s take_screenshot, get_page_info, and capture_pdf tools. Every feature is included on every plan: Free provides 1,000 shots per month with no card; Starter is $5 for 3,000; Growth is $15 for 15,000; Pro is $39 for 60,000; Scale is $99 for 250,000; and Business is $249 for 1,000,000. Yearly billing gives two months free. Sign up free for 1,000 screenshots a month with no card.

FAQ

Does XPath preserve the order of matched elements?

Yes. Iterating the returned node list follows document order, so the sequence corresponds to the order of elements in the parsed tree.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I change the parsed document after selecting nodes?

Yes. The results are DOM nodes belonging to the document, so later DOM edits are visible through those node objects. Re-run a query when you need a newly calculated collection after structural changes.

Should I use CSS or XPath for a selector that users configure?

Validate and constrain user-supplied selectors before evaluating them. For fixed application selectors, either CSS through DomCrawler or XPath through DOMXPath is appropriate; for unrestricted input, add an allowlist and error handling rather than executing arbitrary expressions.

Frequently Asked Questions

Does XPath preserve the order of matched elements?

Yes. Iterating the returned node list follows document order.

Can I change the parsed document after selecting nodes?

Yes. Results are DOM nodes; re-run the query when you need a refreshed collection after structural edits.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use CSS or XPath for a selector that users configure?

Validate and constrain user-supplied selectors with an allowlist and error handling before evaluating them.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.