Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the least powerful layer that can reliably obtain the data. Start with an HTTP crawler such as Colly, parse the returned HTML with goquery, and add a browser controller such as chromedp only when the page requires JavaScript execution or user-like interaction. Colly and goquery are complementary, not competing packages: Colly can fetch and discover pages while goquery extracts fields from each response.

The three layers at a glance

Need Best fit What it does What it does not do
Crawl many ordinary HTTP pages Colly Schedules requests, follows links, runs callbacks, and provides controls for domains, depth, concurrency, cookies, caching and errors. It is not a JavaScript browser or an HTML selector library.
Select data from returned HTML goquery Loads an HTML document and offers chainable, jQuery-like querying and manipulation methods. It does not fetch URLs, execute scripts or click controls.
Render and interact with a browser page chromedp Controls a browser that supports Chrome DevTools Protocol (CDP), including navigation, DOM queries, waits and interaction. It is not a replacement for crawl scheduling, rate limiting or data modeling.

These roles come from the projects’ documented purposes. There is no controlled, directly comparable benchmark establishing a universal speed winner, so measure your own target sites and workload instead of repeating a generic claim.

Choose the layer before writing code

Use Colly plus goquery for server-delivered HTML

Make a normal request first. If the product names, article text, links or metadata are already in the response body, a browser adds deployment and resource overhead without improving the result. Colly handles URL discovery and request coordination; goquery handles selectors and text extraction.

Use chromedp when browser execution is part of the requirement

Choose a browser when content appears only after JavaScript runs, an interaction (such as opening a menu or clicking “load more”) changes the DOM, or you need browser-specific behavior. Expect to manage a Chrome/Chromium executable, contexts, wait conditions and higher operational complexity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a hybrid pipeline for mixed sites

A practical crawler can use Colly for most URLs and send only JavaScript-dependent pages to a chromedp worker. Keep extraction code separate from fetch code so a page can be moved between HTTP and browser paths without rewriting your data model.

Installation and project setup

Create a module and add the packages you need. Package paths and versions change; check the current module metadata and compatibility with your Go release before pinning production builds. The goquery result historically appeared under a legacy gopkg.in/goquery.v1 path, so verify the canonical module path before installing.

mkdir go-scraper && cd go-scraper
go mod init example.com/go-scraper
go get github.com/gocolly/colly/v2
go get github.com/PuerkitoBio/goquery
go get github.com/chromedp/chromedp

The chromedp package listing identified v0.16.0 as published on 2026-07-14; treat that as a dated reference, not a promise that it is the newest release when you build.

A complete Colly and goquery crawler

This example restricts requests to one host, extracts article titles, follows links only within the allowed domain, and reports failures. Replace the URL and selectors with those for your site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
package main

import (
    "fmt"
    "log"
    "strings"
    "time"

    "github.com/PuerkitoBio/goquery"
    "github.com/gocolly/colly/v2"
)

func main() {
    c := colly.NewCollector(
        colly.AllowedDomains("example.com", "www.example.com"),
        colly.MaxDepth(2),
        colly.IgnoreRobotsTxt(false),
    )

    // Keep traffic deliberate for the target host.
    c.Limit(&colly.LimitRule{
        DomainGlob:  "example.com/*",
        Parallelism: 2,
        Delay:       750 * time.Millisecond,
    })

    c.OnHTML("article", func(e *colly.HTMLElement) {
        title := strings.TrimSpace(e.ChildText("h1"))
        summary := strings.TrimSpace(e.ChildText(".summary"))
        fmt.Printf("%qn%snn", title, summary)
    })

    c.OnResponse(func(r *colly.Response) {
        // Parse the complete response with goquery when selectors need more control.
        doc, err := goquery.NewDocumentFromReader(strings.NewReader(string(r.Body)))
        if err != nil {
            log.Printf("parse %s: %v", r.Request.URL, err)
            return
        }
        canonical, ok := doc.Find("link[rel='canonical']").Attr("href")
        if ok {
            log.Printf("canonical: %s", canonical)
        }
    })

    c.OnHTML("a[href]", func(e *colly.HTMLElement) {
        if err := e.Request.Visit(e.Request.AbsoluteURL(e.Attr("href"))); err != nil {
            // ErrAlreadyVisited is normal; log other failures in production code.
            log.Printf("visit %s: %v", e.Request.URL, err)
        }
    })

    c.OnError(func(r *colly.Response, err error) {
        log.Printf("%d %s: %v", r.StatusCode, r.Request.URL, err)
    })

    if err := c.Visit("https://example.com/"); err != nil {
        log.Fatal(err)
    }
}

For a small one-off parse, goquery can work directly on a response body:

doc, err := goquery.NewDocumentFromReader(bytes.NewReader(body))
if err != nil { return err }
doc.Find("article h2").Each(func(_ int, s *goquery.Selection) {
    fmt.Println(strings.TrimSpace(s.Text()))
})

Selectors are only as stable as the site’s markup. Prefer semantic attributes or documented data attributes over deeply nested CSS paths, and validate that a selector found the expected number of elements.

Important Colly controls for safe crawling

Scope and depth

Use allowed domains, URL filters and MaxDepth to prevent accidental expansion into unrelated hosts or infinite calendar/search links. Add request-count limits in your own application when a hard budget matters.

Robots.txt and site rules

Colly checks robots.txt unless configured otherwise. That is an implementation control, not a legal opinion about permission to access a site. Review the target’s terms, robots directives, authentication requirements and applicable law before collecting data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Rate, concurrency and retries

Set per-domain delay and a modest parallelism value, then increase only after observing server responses. Implement bounded retries for transient network errors, while avoiding retries for permanent 4xx responses. Record status codes, elapsed time and the final URL for diagnosis.

Cookies, headers and caching

Use a cookie jar for sessions that the site permits, send only necessary headers, and cache responses when freshness allows. Caching reduces duplicate requests but can hide changes; choose an explicit expiration policy and invalidate it when the source updates.

Browser automation with chromedp

chromedp drives a CDP-compatible browser. The following program starts a headless browser, waits for a rendered selector, reads the resulting DOM text and clicks a button before reading again.

package main

import (
    "context"
    "fmt"
    "log"
    "time"

    "github.com/chromedp/chromedp"
)

func main() {
    ctx, cancel := chromedp.NewContext(context.Background())
    defer cancel()

    ctx, cancel = context.WithTimeout(ctx, 45*time.Second)
    defer cancel()

    var rendered string
    err := chromedp.Run(ctx,
        chromedp.Navigate("https://example.com/app"),
        chromedp.WaitVisible("main", chromedp.ByQuery),
        chromedp.Text("main", &rendered, chromedp.ByQuery),
        chromedp.Click("button.load-more", chromedp.ByQuery),
        chromedp.WaitVisible(".results .item", chromedp.ByQuery),
        chromedp.Text("main", &rendered, chromedp.ByQuery),
    )
    if err != nil {
        log.Fatal(err)
    }
    fmt.Println(rendered)
}

In deployment, make the browser executable available, set a timeout for every navigation, and wait for a meaningful application state rather than an arbitrary sleep. A network-idle wait can still fire too early on pages with polling; a specific selector or application-ready marker is usually more deterministic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

For a single clean screenshot or PDF, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

One request returns PNG, JPEG, WebP or PDF. The API supports full-page and element captures, device presets and custom viewports, dark mode, retina scale, PDF paper and page-range settings, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which eases migration.

For developers who want AI-assisted browser work, its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for request options. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability and cost decisions

  • HTTP first: Colly and goquery avoid browser startup and generally use fewer resources, but no reviewed source supplies a controlled comparison against browser extraction.
  • Browser selectively: Route only JavaScript-dependent URLs to chromedp and reuse browser contexts where safe. Bound tabs, memory and navigation timeouts.
  • Measure the workload: Track pages per minute, error rate, median and tail latency, bytes transferred and extraction completeness on representative URLs.
  • Make results reproducible: Store source URL, retrieval time, status, selector version and parser errors. HTML changes are an extraction failure, not necessarily a network failure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“The selector returns nothing”

Inspect the raw HTTP body. If the data is absent, switch to chromedp or locate the underlying JSON endpoint that the site is permitted to expose. If it is present, check namespaces, encoding and changed class names.

“The browser times out”

Confirm Chrome/Chromium is installed and executable, extend the context timeout cautiously, and wait for a stable selector instead of a fixed delay. Capture console and network errors when diagnosing an application failure.

“The crawler visits too many URLs”

Add allowed domains, URL regular expressions, maximum depth and an application-level request budget. Exclude tracking parameters and unbounded search or calendar paths.

“Requests are rejected”

Slow the per-domain rate, honor robots.txt and access rules, use an appropriate user agent, and stop on persistent 403, 429 or CAPTCHA responses. Do not treat a technical workaround as permission.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Results are stale or duplicated”

Review cache TTL and canonicalization, normalize URLs before visiting, and record a content hash or last-seen timestamp so updates can be distinguished from repeats.

A practical decision checklist

  1. Fetch one target URL and inspect the returned HTML.
  2. If the required fields are present, use Colly for discovery and goquery for extraction.
  3. If fields appear only after script execution or interaction, prototype that path in chromedp.
  4. Define domain, depth, URL, request-count and rate limits before scaling.
  5. Add timeouts, bounded retries, caching and structured error logging.
  6. Benchmark representative pages yourself; do not infer a speed ranking from package descriptions.
  7. Recheck module versions, browser compatibility and site rules at deployment time.

FAQ

Do Colly and goquery replace each other?

No. Colly coordinates requests and crawling; goquery queries an HTML document. A common design uses both.

Do I need a browser to scrape a JavaScript-rendered page?

Only when the required content or interaction is unavailable in the initial HTML or an accessible data response. Verify the response before adding browser automation.

Is chromedp a cross-browser framework?

It controls browsers that support Chrome DevTools Protocol. The available evidence does not establish a current cross-browser Go framework comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I ignore robots.txt with Colly?

Colly exposes a configuration switch, but disabling a check does not grant permission. Follow the target site’s rules and applicable law.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.