Free tools Windows power users keep installed
One-click scans. No signup required.
Use the least powerful layer that can reliably obtain the data. Start with an HTTP crawler such as Colly, parse the returned HTML with goquery, and add a browser controller such as chromedp only when the page requires JavaScript execution or user-like interaction. Colly and goquery are complementary, not competing packages: Colly can fetch and discover pages while goquery extracts fields from each response.
The three layers at a glance
| Need | Best fit | What it does | What it does not do |
|---|---|---|---|
| Crawl many ordinary HTTP pages | Colly | Schedules requests, follows links, runs callbacks, and provides controls for domains, depth, concurrency, cookies, caching and errors. | It is not a JavaScript browser or an HTML selector library. |
| Select data from returned HTML | goquery | Loads an HTML document and offers chainable, jQuery-like querying and manipulation methods. | It does not fetch URLs, execute scripts or click controls. |
| Render and interact with a browser page | chromedp | Controls a browser that supports Chrome DevTools Protocol (CDP), including navigation, DOM queries, waits and interaction. | It is not a replacement for crawl scheduling, rate limiting or data modeling. |
These roles come from the projects’ documented purposes. There is no controlled, directly comparable benchmark establishing a universal speed winner, so measure your own target sites and workload instead of repeating a generic claim.
Choose the layer before writing code
Use Colly plus goquery for server-delivered HTML
Make a normal request first. If the product names, article text, links or metadata are already in the response body, a browser adds deployment and resource overhead without improving the result. Colly handles URL discovery and request coordination; goquery handles selectors and text extraction.
Use chromedp when browser execution is part of the requirement
Choose a browser when content appears only after JavaScript runs, an interaction (such as opening a menu or clicking “load more”) changes the DOM, or you need browser-specific behavior. Expect to manage a Chrome/Chromium executable, contexts, wait conditions and higher operational complexity.
#1 Best Overall
Use a hybrid pipeline for mixed sites
A practical crawler can use Colly for most URLs and send only JavaScript-dependent pages to a chromedp worker. Keep extraction code separate from fetch code so a page can be moved between HTTP and browser paths without rewriting your data model.
Installation and project setup
Create a module and add the packages you need. Package paths and versions change; check the current module metadata and compatibility with your Go release before pinning production builds. The goquery result historically appeared under a legacy gopkg.in/goquery.v1 path, so verify the canonical module path before installing.
mkdir go-scraper && cd go-scraper
go mod init example.com/go-scraper
go get github.com/gocolly/colly/v2
go get github.com/PuerkitoBio/goquery
go get github.com/chromedp/chromedp
The chromedp package listing identified v0.16.0 as published on 2026-07-14; treat that as a dated reference, not a promise that it is the newest release when you build.
A complete Colly and goquery crawler
This example restricts requests to one host, extracts article titles, follows links only within the allowed domain, and reports failures. Replace the URL and selectors with those for your site.
package main
import (
"fmt"
"log"
"strings"
"time"
"github.com/PuerkitoBio/goquery"
"github.com/gocolly/colly/v2"
)
func main() {
c := colly.NewCollector(
colly.AllowedDomains("example.com", "www.example.com"),
colly.MaxDepth(2),
colly.IgnoreRobotsTxt(false),
)
// Keep traffic deliberate for the target host.
c.Limit(&colly.LimitRule{
DomainGlob: "example.com/*",
Parallelism: 2,
Delay: 750 * time.Millisecond,
})
c.OnHTML("article", func(e *colly.HTMLElement) {
title := strings.TrimSpace(e.ChildText("h1"))
summary := strings.TrimSpace(e.ChildText(".summary"))
fmt.Printf("%qn%snn", title, summary)
})
c.OnResponse(func(r *colly.Response) {
// Parse the complete response with goquery when selectors need more control.
doc, err := goquery.NewDocumentFromReader(strings.NewReader(string(r.Body)))
if err != nil {
log.Printf("parse %s: %v", r.Request.URL, err)
return
}
canonical, ok := doc.Find("link[rel='canonical']").Attr("href")
if ok {
log.Printf("canonical: %s", canonical)
}
})
c.OnHTML("a[href]", func(e *colly.HTMLElement) {
if err := e.Request.Visit(e.Request.AbsoluteURL(e.Attr("href"))); err != nil {
// ErrAlreadyVisited is normal; log other failures in production code.
log.Printf("visit %s: %v", e.Request.URL, err)
}
})
c.OnError(func(r *colly.Response, err error) {
log.Printf("%d %s: %v", r.StatusCode, r.Request.URL, err)
})
if err := c.Visit("https://example.com/"); err != nil {
log.Fatal(err)
}
}
For a small one-off parse, goquery can work directly on a response body:
doc, err := goquery.NewDocumentFromReader(bytes.NewReader(body))
if err != nil { return err }
doc.Find("article h2").Each(func(_ int, s *goquery.Selection) {
fmt.Println(strings.TrimSpace(s.Text()))
})
Selectors are only as stable as the site’s markup. Prefer semantic attributes or documented data attributes over deeply nested CSS paths, and validate that a selector found the expected number of elements.
Important Colly controls for safe crawling
Scope and depth
Use allowed domains, URL filters and MaxDepth to prevent accidental expansion into unrelated hosts or infinite calendar/search links. Add request-count limits in your own application when a hard budget matters.
Robots.txt and site rules
Colly checks robots.txt unless configured otherwise. That is an implementation control, not a legal opinion about permission to access a site. Review the target’s terms, robots directives, authentication requirements and applicable law before collecting data.
Rate, concurrency and retries
Set per-domain delay and a modest parallelism value, then increase only after observing server responses. Implement bounded retries for transient network errors, while avoiding retries for permanent 4xx responses. Record status codes, elapsed time and the final URL for diagnosis.
Cookies, headers and caching
Use a cookie jar for sessions that the site permits, send only necessary headers, and cache responses when freshness allows. Caching reduces duplicate requests but can hide changes; choose an explicit expiration policy and invalidate it when the source updates.
Browser automation with chromedp
chromedp drives a CDP-compatible browser. The following program starts a headless browser, waits for a rendered selector, reads the resulting DOM text and clicks a button before reading again.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/chromedp/chromedp"
)
func main() {
ctx, cancel := chromedp.NewContext(context.Background())
defer cancel()
ctx, cancel = context.WithTimeout(ctx, 45*time.Second)
defer cancel()
var rendered string
err := chromedp.Run(ctx,
chromedp.Navigate("https://example.com/app"),
chromedp.WaitVisible("main", chromedp.ByQuery),
chromedp.Text("main", &rendered, chromedp.ByQuery),
chromedp.Click("button.load-more", chromedp.ByQuery),
chromedp.WaitVisible(".results .item", chromedp.ByQuery),
chromedp.Text("main", &rendered, chromedp.ByQuery),
)
if err != nil {
log.Fatal(err)
}
fmt.Println(rendered)
}
In deployment, make the browser executable available, set a timeout for every navigation, and wait for a meaningful application state rather than an arbitrary sleep. A network-idle wait can still fire too early on pages with polling; a specific selector or application-ready marker is usually more deterministic.
Recommended Free Tools
Or skip the browser setup
For a single clean screenshot or PDF, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
One request returns PNG, JPEG, WebP or PDF. The API supports full-page and element captures, device presets and custom viewports, dark mode, retina scale, PDF paper and page-range settings, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which eases migration.
For developers who want AI-assisted browser work, its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for request options. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.
Rank #4
Performance, reliability and cost decisions
- HTTP first: Colly and goquery avoid browser startup and generally use fewer resources, but no reviewed source supplies a controlled comparison against browser extraction.
- Browser selectively: Route only JavaScript-dependent URLs to chromedp and reuse browser contexts where safe. Bound tabs, memory and navigation timeouts.
- Measure the workload: Track pages per minute, error rate, median and tail latency, bytes transferred and extraction completeness on representative URLs.
- Make results reproducible: Store source URL, retrieval time, status, selector version and parser errors. HTML changes are an extraction failure, not necessarily a network failure.
Troubleshooting common failures
“The selector returns nothing”
Inspect the raw HTTP body. If the data is absent, switch to chromedp or locate the underlying JSON endpoint that the site is permitted to expose. If it is present, check namespaces, encoding and changed class names.
“The browser times out”
Confirm Chrome/Chromium is installed and executable, extend the context timeout cautiously, and wait for a stable selector instead of a fixed delay. Capture console and network errors when diagnosing an application failure.
“The crawler visits too many URLs”
Add allowed domains, URL regular expressions, maximum depth and an application-level request budget. Exclude tracking parameters and unbounded search or calendar paths.
“Requests are rejected”
Slow the per-domain rate, honor robots.txt and access rules, use an appropriate user agent, and stop on persistent 403, 429 or CAPTCHA responses. Do not treat a technical workaround as permission.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
“Results are stale or duplicated”
Review cache TTL and canonicalization, normalize URLs before visiting, and record a content hash or last-seen timestamp so updates can be distinguished from repeats.
Best Value
A practical decision checklist
- Fetch one target URL and inspect the returned HTML.
- If the required fields are present, use Colly for discovery and goquery for extraction.
- If fields appear only after script execution or interaction, prototype that path in chromedp.
- Define domain, depth, URL, request-count and rate limits before scaling.
- Add timeouts, bounded retries, caching and structured error logging.
- Benchmark representative pages yourself; do not infer a speed ranking from package descriptions.
- Recheck module versions, browser compatibility and site rules at deployment time.
FAQ
Do Colly and goquery replace each other?
No. Colly coordinates requests and crawling; goquery queries an HTML document. A common design uses both.
Do I need a browser to scrape a JavaScript-rendered page?
Only when the required content or interaction is unavailable in the initial HTML or an accessible data response. Verify the response before adding browser automation.
Is chromedp a cross-browser framework?
It controls browsers that support Chrome DevTools Protocol. The available evidence does not establish a current cross-browser Go framework comparison.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Can I ignore robots.txt with Colly?
Colly exposes a configuration switch, but disabling a check does not grant permission. Follow the target site’s rules and applicable law.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

