Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →For straightforward OCaml web scraping, use Cohttp to make the HTTP request and Lambda Soup to parse the returned HTML and select content with CSS selectors. Choose Cohttp’s backend to match your application’s runtime—Lwt, Async, curl, or Eio. If you need streaming parsing or lower-level control, consider Markup.ml. These libraries fetch and parse HTML; the available documentation does not establish that they run a site’s JavaScript.
How an OCaml scraper fits together
Scraping a page has two distinct stages: retrieving its response and interpreting the HTML in that response. Cohttp provides HTTP client implementations; Lambda Soup and Markup.ml handle HTML parsing and extraction. Keeping the jobs separate makes it easier to diagnose whether a problem comes from the network response or from your selectors.
- Request: use a Cohttp client backend appropriate to the concurrency runtime and deployment environment.
- Inspect the response: check the status and response headers, and decide how your application should handle unsuccessful responses.
- Parse: give the returned HTML to Lambda Soup for a document-oriented workflow, or use Markup.ml when streaming or parser-signal control matters.
- Extract: use selectors and traversals to collect the required text or attributes.
- Validate: test the selectors against the actual pages you intend to process, including pages whose structure or content varies.
A successful HTTP request does not guarantee that the response contains the content you see in a browser. A page may require client-side JavaScript, return an access challenge, or present different HTML to automated clients. Treat those as properties to investigate for each target, not capabilities guaranteed by the OCaml libraries.
Choose the libraries and runtime
| Need | Tool | What it provides | When to choose it |
|---|---|---|---|
| HTTP requests | Cohttp and a backend package | HTTP client implementations for Lwt, Async, curl, and Eio | Match the backend to the concurrency runtime and deployment target already used by your application. |
| Document-oriented HTML extraction | Lambda Soup | CSS selectors, lazy traversals, text extraction, and DOM mutation | Choose it when selecting elements from a parsed document is the natural way to express the extraction. |
| Streaming or lower-level parsing | Markup.ml | HTML5 and XML parsing, lazy signal streams, single-pass streaming, and error recovery | Consider it for large or streaming inputs, or when you need direct control of parser signals. |
| HTML generation | TyXML | Typed combinators for producing valid HTML and SVG | This is adjacent web tooling, not a scraping or HTML-extraction library. |
Lambda Soup’s documentation describes it as an HTML scraping library inspired by Python’s Beautiful Soup. It is based on Markup.ml. That makes the two useful at different levels: Lambda Soup offers a convenient document and selector interface, while Markup.ml exposes parsing primitives suitable for stream-oriented work.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
Pick the Cohttp backend deliberately
Cohttp has separate backend packages rather than one universal runtime choice. The documented options include Lwt, Async, curl, and Eio. Use the option that fits your application’s existing event loop and deployment requirements instead of introducing a second concurrency model solely for one scraper.
As package-catalog observations for August 2026, Cohttp and Cohttp Eio were listed at version 6.3.0. The Eio package describes direct-style programming and multicore support for OCaml 5.0 and newer. These listings are not compatibility guarantees: check current package constraints and your OCaml version before pinning dependencies.
Check package constraints before installation
Lambda Soup’s package listing showed version 1.1.1, published September 5, 2024, and Markup.ml’s listing showed version 1.0.3. Package listings change over time, and these version observations do not establish that every combination of Cohttp backend, parser, and compiler is compatible. Review the opam constraints for the versions you plan to use, particularly when adding a backend to an existing project.
Rank #2
Build a fetch-and-extract workflow
The package documentation supports this division of work: make the request with a Cohttp client, then parse the returned HTML with Lambda Soup. The following is a workflow outline, not a verified, copy-and-run program: select the client module and response-body handling appropriate to your installed Cohttp backend, then use Lambda Soup’s documented parser and CSS-selector approach.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Install the needed packages. Add the Cohttp backend matching your runtime and Lambda Soup through opam. The backend is a separate choice; do not assume installing a Cohttp package automatically gives you every runtime implementation.
- Issue an HTTP request. Supply the target URL to the selected client. Handle the response using that backend’s API, and retain the status and headers as well as the body.
- Decide whether the response is usable. Treat status codes and content type as input to your application’s error handling. A body can be an error page or challenge rather than the page you expected.
- Parse the HTML string. Lambda Soup documents parsing an HTML string and selecting elements with CSS selectors. Start with selectors tied to stable page structure rather than assuming every page uses the same markup.
- Extract only the fields you need. Use the selected elements’ text for visible text, or retrieve the relevant attribute when the data is carried in an attribute such as a link destination.
- Check the output. Test representative pages, including missing elements and malformed or incomplete HTML. Treat a missing match as a normal case your scraper handles, not an assumption that parsing always succeeds in the desired shape.
For the parser portion, Lambda Soup’s documented pattern is to parse an HTML string and apply a CSS selector to the resulting document. The exact client call and body conversion depend on whether you chose Lwt, Async, curl, or Eio; consult that backend’s current package documentation rather than copying an API example for a different runtime.
When to use Markup.ml directly
Use Lambda Soup when your extraction is naturally expressed as “find these elements, then read their text or attributes.” Consider Markup.ml directly when you need its lazy signal stream, want a single-pass streaming process, or need more control over parsing. Its package documentation describes both HTML and XML parsers, with error recovery.
Rank #3
Streaming can be useful when the input or processing model makes retaining a full document inconvenient. It does not, by itself, solve HTTP fetching, JavaScript rendering, or site access restrictions. Those remain separate concerns.
JavaScript-rendered pages and browser-based capture
The documented Cohttp, Lambda Soup, and Markup.ml capabilities establish HTTP clients and HTML parsing; they do not establish browser execution or JavaScript rendering. If the data is absent from the HTTP response because a page builds it client-side, a request-plus-parser workflow may not contain the information your selector needs. Inspect the received HTML before changing selectors. If the content is not present, investigate whether the site exposes an appropriate response or whether a browser-based approach is necessary.
For browser screenshots rather than custom OCaml extraction, ScreenshotNeo offers a website screenshot API and MCP server. It can capture a rendered page as PNG, JPEG, WebP, or PDF; it is an alternative capture workflow, not a replacement for a scraper that needs structured data fields.
Rank #4
- Used Book in Good Condition
Or skip the browser setup
If your task is to capture a page rather than build an OCaml extraction pipeline, one GET request to ScreenshotNeo returns a screenshot or PDF. See the ScreenshotNeo API documentation for available options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses include headers identifying the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteReliability, performance, and responsible use
Do not infer speed from library features
The package descriptions do not provide a comparative throughput benchmark for Cohttp, Lambda Soup, or Markup.ml. Choose the runtime and parser based on your application’s needs, then measure your own workload if performance is important. Page size, network conditions, server behavior, and extraction work all affect end-to-end time; the available package information does not support a claim that one of these libraries is universally faster.
Best Value
Make failures visible in your application
- Record the requested URL and response status so that an HTTP failure is distinguishable from a selector that found nothing.
- Check response headers and body shape before assuming the result is the intended page.
- Handle missing elements and changing markup explicitly; page structure is target-specific.
- For large or streaming inputs, assess whether Markup.ml’s streaming interface fits better than a document-oriented parse.
- Confirm the target’s terms and applicable rules before automating access. A library’s technical capabilities do not determine whether scraping a particular site is permitted.
Troubleshooting common problems
| Symptom | Likely explanation | What to check |
|---|---|---|
| The expected selector returns no elements | The response markup differs from your assumption, the selector is wrong, or the content is not in the server-returned HTML. | Inspect the response body first. Confirm the element and selector against that HTML; if the content is built client-side, a parser alone will not render it. |
| The parser returns content, but it is an error or challenge page | The server returned a different page than the one a normal browser displays. | Check status, headers, and body before extraction. Do not treat a syntactically parseable document as proof that the intended page loaded. |
| Examples for a different client do not compile | Cohttp offers separate backend packages and APIs, so code for one runtime may not match another. | Identify the installed backend and use its current client and response-body interface. |
| Large input is awkward to process as a document | A whole-document workflow may not fit the input or memory constraints. | Evaluate Markup.ml’s lazy, single-pass streaming API and parser signals. |
| Dependency resolution fails | The selected package versions or compiler constraints may not agree. | Review the opam constraints for the Cohttp backend, Lambda Soup, Markup.ml, and your OCaml version before pinning the dependency set. |
Choosing a practical starting point
For a typical scraper that receives usable HTML in an HTTP response, start with a Cohttp backend already compatible with your application and Lambda Soup for CSS-based extraction. Choose Markup.ml directly when streaming or parser-level control is a real requirement. Before expanding the implementation, verify what the target actually returns, confirm the extracted fields across representative pages, and check the site’s rules for automated access.
Frequently Asked Questions
Does OCaml have an HTML scraping library?
Yes. Lambda Soup provides CSS selectors and document traversals for HTML extraction; Cohttp handles HTTP requests, so they cover separate parts of a scraper.
Can Lambda Soup scrape XML?
Markup.ml documents HTML and XML parsers, while Lambda Soup is described as an HTML scraping library. For XML parsing needs, evaluate Markup.ml directly.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

