Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For pages whose needed content is already in the HTML response, use HttpClient to fetch the page and a parser such as Html Agility Pack or AngleSharp to extract data. If the content appears only after JavaScript runs or requires browser interaction, use Playwright for .NET. The production-ready part is less about choosing a fashionable library and more about managing connections, respecting access constraints, validating changing markup, and making failures visible.

How do I scrape a website with C#?

Start by checking whether the page’s initial HTTP response contains the information you need. If it does, a lightweight pipeline is usually sufficient:

  1. Send an asynchronous HTTP request with a reusable HttpClient.
  2. Check the response status and read the HTML.
  3. Parse the document with Html Agility Pack or AngleSharp.
  4. Extract and normalize the fields you need, allowing for missing or changed elements.
  5. Validate the results and persist them in your chosen store.

Use browser automation only when the page depends on JavaScript execution, browser state, or interactions that a direct request cannot reproduce. This distinction avoids running a browser for pages that can be handled with ordinary HTTP.

A minimal static-page example

The following console example uses Html Agility Pack and .NET’s built-in dependency injection abstractions to configure an HTTP client. Add the packages HtmlAgilityPack and Microsoft.Extensions.Http to a .NET console project, then use this as Program.cs. Replace the example URL and selectors with ones for a site you are authorized to access.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
using HtmlAgilityPack;
using Microsoft.Extensions.DependencyInjection;

var services = new ServiceCollection();
services.AddHttpClient("scraper", client =>
{
    client.DefaultRequestHeaders.UserAgent.ParseAdd("ExampleResearchBot/1.0");
});

using var provider = services.BuildServiceProvider();
var factory = provider.GetRequiredService<IHttpClientFactory>();
var client = factory.CreateClient("scraper");

var url = "https://example.com/";
using var response = await client.GetAsync(url);
response.EnsureSuccessStatusCode();

var html = await response.Content.ReadAsStringAsync();
var document = new HtmlDocument();
document.LoadHtml(html);

var titleNode = document.DocumentNode.SelectSingleNode("//title");
var title = HtmlEntity.DeEntitize(titleNode?.InnerText ?? "").Trim();

var links = document.DocumentNode.SelectNodes("//a[@href]") ??
            new HtmlNodeCollection(null);
foreach (var link in links)
{
    var href = link.GetAttributeValue("href", "");
    var text = HtmlEntity.DeEntitize(link.InnerText).Trim();
    if (href.Length > 0)
        Console.WriteLine($"{text}t{href}");
}

Console.WriteLine($"Page title: {title}");

This example deliberately handles a missing title and an absent link collection instead of assuming the page matches a fixed template. For real data, also resolve relative links against the page URI, normalize whitespace, validate required values, and record when expected fields are missing. Persisting an empty string as if it were valid data can conceal a site redesign or selector mistake.

Which C# library should I use for web scraping?

Choose the tool based on where the content lives and what operations the page requires. Html Agility Pack and AngleSharp parse HTML; neither runs page JavaScript. Playwright controls an actual browser and is appropriate when rendering or interaction is necessary.

Tool Use it for Trade-off
HttpClient Making HTTP requests and receiving responses Does not execute JavaScript or provide a browser DOM.
Html Agility Pack Parsing returned HTML, often with XPath selection Does not render a page; selector and document-handling choices should fit the markup and team.
AngleSharp Parsing HTML with a standards-oriented DOM and CSS-selector style APIs Does not execute page JavaScript as a browser would.
Playwright for .NET Browser rendering, interactions, and observing browser requests and responses Requires browser binaries and adds runtime and deployment overhead.

There is no established benchmark here showing one parser is objectively faster or better for every document. Choose between Html Agility Pack and AngleSharp by trying the selectors your target markup needs, the document behavior you encounter, API ergonomics, and team familiarity. Microsoft’s ASP.NET Core integration-testing documentation mentions both parsers in its sample context; it is not a comparative performance study or a current scraping recommendation.

Should I use HttpClient, HtmlAgilityPack, AngleSharp, or Playwright?

These are not four interchangeable scrapers. HttpClient fetches a response; a parser interprets returned HTML; Playwright runs a browser. A common static-page combination is HttpClient plus one parser. A dynamic-page combination may use Playwright to render and interact, with browser locators or DOM evaluation to extract what is visible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use HttpClient with a parser when

  • The response body contains the content you need.
  • You can access the relevant page without browser-only actions.
  • You want to avoid installing and operating browser binaries.

Use Playwright when

  • The content is inserted after scripts execute.
  • The page requires interactions, such as clicking a control or navigating through a browser flow.
  • Browser request and response events would help you understand how the page obtains its data.

Playwright’s .NET port automates Chromium, Firefox, and WebKit. Browser request events can expose useful details about the page’s network activity. Keep in mind that a request can complete even when the HTTP response is an error such as 404 or 503; inspect response status as well as whether a request completed.

Can C# scrape JavaScript-rendered pages?

Yes, when you use a browser automation tool such as Playwright for .NET. A normal HTTP client receives the server’s response; a parser can select elements from that response, but it cannot execute the page’s JavaScript. If a required value is missing from the response HTML and appears only after script execution, parsing the original HTML will not make that value appear.

Before adding browser automation, compare the initial response with the rendered page. Browser network events may reveal that the page retrieves data through a request; if an appropriate public endpoint exists and its use is permitted, direct HTTP access to that endpoint may be simpler than rendering the entire page. Do not assume that discovering an endpoint grants permission to use it.

For a browser-based workflow, install the Playwright .NET package, install the browser binaries required by your chosen browser, launch the browser, navigate to the permitted page, wait for the specific content you need, and extract it. Use a condition tied to the target element or state rather than an arbitrary sleep where possible. Browser installation, browser versioning, and deployment environment are additional operational responsibilities compared with an HTTP-and-parser workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to make a scraper reliable in production

Manage HttpClient lifetime and DNS changes

Do not create and dispose a new HttpClient for every request. Microsoft recommends either long-lived clients configured with PooledConnectionLifetime or short-lived clients created by IHttpClientFactory. Both approaches reuse connection pools rather than repeatedly creating clients without regard to connection lifetime.

DNS is resolved when a connection is created. A long-lived connection can therefore continue using an endpoint associated with an earlier DNS lookup; HttpClient does not continually apply DNS TTL changes to an already-open connection. PooledConnectionLifetime allows connections to be replaced so DNS can be resolved again. Microsoft’s 15-minute sample is illustrative, not a universal production setting. Choose a lifetime based on the service and expected DNS changes, or use the factory pattern where it suits the application’s configuration needs.

Handle cancellation, timeouts, and response failures

Use cancellation tokens through the request and any downstream work so an application can stop a scrape that is no longer needed. Configure timeouts for your workload rather than copying an unexplained magic number. Check HTTP status before treating a response body as a successful page, and distinguish network failures, timeouts, access denials, and error responses in logs. A retry is not automatically safe: retry only when the failure and operation make another request appropriate.

Set operational bounds around response size and concurrency. A very large body can consume unnecessary memory; a burst of parallel requests can burden a target site or trigger its controls. Use bounded concurrency and site-appropriate pacing. There is no universal safe request rate or concurrency number: the target’s instructions, capacity, and your authorization matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make extraction failures observable

Web pages change. Treat expected fields as data with validation rules: check required values, normalize whitespace and entities, and count or record missing fields. Preserve enough context to diagnose a failure, such as the URL, response status, timestamp, and which expected selector was absent, while avoiding unnecessary collection of personal information. Monitor for sudden changes in empty results or parse failures instead of silently writing corrupted records.

Keep browser work bounded

When using Playwright, close pages, contexts, and browser processes deterministically. Limit how many pages run concurrently according to the available resources and the target site’s capacity. Browser automation consumes more operational resources than a direct HTTP request and should be reserved for requirements such as script rendering or interaction.

Is robots.txt permission to scrape?

No. The IETF’s RFC 9309 says, “These rules are not a form of access authorization.” A robots.txt file communicates crawler instructions; it does not grant permission, override access controls, or settle whether a particular use is lawful.

RFC 9309 describes how crawlers that honor the Robots Exclusion Protocol should handle the file. When it is successfully retrieved, parseable rules are to be followed. The standard says cached robots.txt files generally should not be used for more than 24 hours unless the file is unreachable. If a server or network error makes the file unreachable, the crawler must assume complete disallow. A 4xx “unavailable” response can be treated differently under the protocol, so “no file found means scraping is allowed” is not a sound general rule.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The standard specifies a minimum parsing limit of 500 kibibytes (KiB). It also gives 30 days as an example duration after which an undefined robots.txt may be treated as unavailable or a cached copy may continue to be used. These are protocol details, not request-rate recommendations or legal deadlines.

For a real target, check its terms, obtain authorization where needed, and consider applicable jurisdiction, access controls, copyright, and privacy obligations. Public accessibility by itself does not establish that a planned scraping use is lawful.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common C# scraping problems and fixes

Symptom Likely cause What to check or change
Expected elements are missing Content is generated by JavaScript, the selector no longer matches, or the server returned an unexpected page. Inspect the raw response and status first. If the content appears only after rendering, use Playwright; otherwise update and validate selectors.
Requests fail intermittently or hang Network, server, timeout, or cancellation behavior is not being handled explicitly. Log the failure category and status, configure a workload-appropriate timeout, propagate cancellation, and retry only when another attempt is appropriate.
Scraper keeps using an old host endpoint A pooled connection remains open after DNS changes. Use a suitable PooledConnectionLifetime for a long-lived client or use IHttpClientFactory.
Results suddenly become blank or malformed The page structure changed, an expected element is missing, or an error page was parsed as content. Validate required fields, check response status, record missing selectors, and alert on a rise in parse failures.
Browser sees a failed resource despite a completed request event The request completed with an HTTP error response. Inspect the response status; a completed browser request does not imply a successful HTTP status.
Target blocks or slows the scraper Request pattern, access restrictions, or target capacity may be involved. Stop and review site instructions and authorization; reduce or bound activity rather than attempting to bypass controls.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo offers a website screenshot API and MCP server for developers. A GET request returns a PNG, JPEG, WebP, or PDF. Here is the one-call cURL example; see the ScreenshotNeo API documentation for its request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for 1,000 free screenshots a month with no card.

Further reading

Frequently Asked Questions

Does Microsoft recommend WebClient for new C# scrapers?

No. Microsoft documents WebRequest, WebClient, and ServicePoint as obsolete beginning with .NET 6 and recommends HttpClient instead.

Do I need a browser just to parse HTML with CSS selectors?

No. AngleSharp provides DOM and CSS-selector-style parsing for returned HTML; browser automation is needed only when script execution or browser interaction is required.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.