Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Go’s standard library gives you the core pieces for a small web scraper: HTTP requests with net/http, URL parsing with net/url, cancellation with context, and response-body streaming with io. For practical HTML5 parsing, you’ll also need golang.org/x/net/html, a separately versioned module—not part of the standard library.

The example below fetches one page, limits the response size, parses its HTML, and prints links. It is a starting point, not a full crawler: add deliberate scope, rate, and site-policy controls before following links at scale.

What Go packages does a basic scraper need?

Task Package How it fits
Send HTTP requests net/http Use a reusable client to fetch pages and control request behavior.
Parse and resolve URLs net/url Validate target URLs, build query parameters, and resolve relative links.
Stop work on cancellation or deadline context Attach a context to each request so it can be canceled or time-limited.
Read response data io Treat the response body as a stream and impose a byte limit where needed.
Tokenize or parse HTML golang.org/x/net/html Provides HTML5 tokenization and tree construction; it is an external module.

See the net/http documentation, the Go Authors’ net/http package documentation, and the documentation for net/url, context, io, and golang.org/x/net/html.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a small scraper

This example accepts a URL, requests the page with a timeout, reads at most 2 MiB, parses the HTML tree, and prints each link’s text and resolved URL. The 2 MiB cap is an example policy for this program, not a Go or web standard.

  1. Create a module and add the external HTML package:

    go mod init example.com/scraper
    go get golang.org/x/net/html
  2. Save the following as main.go:

    package main
    
    import (
    	"context"
    	"fmt"
    	"io"
    	"net/http"
    	"net/url"
    	"os"
    	"strings"
    	"time"
    
    	"golang.org/x/net/html"
    )
    
    const maxBody = 2 << 20 // 2 MiB
    
    func main() {
    	if len(os.Args) != 2 {
    		fmt.Fprintln(os.Stderr, "usage: scraper https://example.com/")
    		os.Exit(2)
    	}
    
    	pageURL, err := url.Parse(os.Args[1])
    	if err != nil || pageURL.Host == "" || (pageURL.Scheme != "http" && pageURL.Scheme != "https") {
    		fmt.Fprintln(os.Stderr, "provide a valid http or https URL")
    		os.Exit(2)
    	}
    
    	client := &http.Client{Timeout: 15 * time.Second}
    	ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
    	defer cancel()
    
    	req, err := http.NewRequestWithContext(ctx, http.MethodGet, pageURL.String(), nil)
    	if err != nil {
    		fmt.Fprintln(os.Stderr, "create request:", err)
    		os.Exit(1)
    	}
    
    	resp, err := client.Do(req)
    	if err != nil {
    		fmt.Fprintln(os.Stderr, "fetch page:", err)
    		os.Exit(1)
    	}
    	defer resp.Body.Close()
    
    	if resp.StatusCode < 200 || resp.StatusCode >= 300 {
    		fmt.Fprintf(os.Stderr, "unexpected HTTP status: %sn", resp.Status)
    		os.Exit(1)
    	}
    
    	body, err := io.ReadAll(io.LimitReader(resp.Body, maxBody+1))
    	if err != nil {
    		fmt.Fprintln(os.Stderr, "read response:", err)
    		os.Exit(1)
    	}
    	if len(body) > maxBody {
    		fmt.Fprintf(os.Stderr, "response exceeds %d bytesn", maxBody)
    		os.Exit(1)
    	}
    
    	doc, err := html.Parse(strings.NewReader(string(body)))
    	if err != nil {
    		fmt.Fprintln(os.Stderr, "parse HTML:", err)
    		os.Exit(1)
    	}
    
    	var walk func(*html.Node)
    	walk = func(n *html.Node) {
    		if n.Type == html.ElementNode && n.Data == "a" {
    			for _, attr := range n.Attr {
    				if attr.Key == "href" {
    					ref, err := url.Parse(attr.Val)
    					if err == nil {
    						fmt.Printf("%st%sn", linkText(n), pageURL.ResolveReference(ref))
    					}
    					break
    				}
    			}
    		}
    		for child := n.FirstChild; child != nil; child = child.NextSibling {
    			walk(child)
    		}
    	}
    	walk(doc)
    }
    
    func linkText(n *html.Node) string {
    	var b strings.Builder
    	var walk func(*html.Node)
    	walk = func(node *html.Node) {
    		if node.Type == html.TextNode {
    			b.WriteString(node.Data)
    		}
    		for child := node.FirstChild; child != nil; child = child.NextSibling {
    			walk(child)
    		}
    	}
    	walk(n)
    	return strings.TrimSpace(b.String())
    }

    The example uses both a client timeout and a request context deadline. This makes the time limit explicit at the client and request levels; in a larger application, choose a consistent timeout policy that fits its request lifecycle.

  3. Run it with a page you are allowed to fetch:

    go run . https://example.com/

    Each output line contains the link text followed by its resolved URL. The sample checks the HTTP status before parsing and closes the response body when finished.

Choose a fetching approach

Approach Useful when Trade-off
Convenience helper such as http.Get You need a quick request with default behavior. Less room to express request-specific headers, context, and policy directly.
Explicit request with a reusable http.Client You need headers, cancellation, conditional requests, redirect rules, or client and transport configuration. Requires a little more setup, but makes request behavior visible and reusable.

The Go Authors’ net/http documentation states: “Clients and Transports are safe for concurrent use by multiple goroutines and for efficiency should only be created once and re-used.” Reuse helps avoid unnecessary setup, but it does not mean a scraper should send unlimited concurrent requests to a site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use URL helpers instead of string concatenation

Parse and validate the starting address with url.Parse, use url.Values to encode query parameters, and resolve extracted relative links with ResolveReference. Concatenating URL strings can produce malformed paths, queries, or escaped characters. The net/url documentation describes these parsing and resolution tools.

Choose between an HTML tokenizer and a tree parser

Option Best fit What to consider
HTML tokenizer Extracting data that can be handled token by token without navigating a reconstructed document tree. It exposes lower-level token data; manage token and byte-slice lifetimes carefully.
HTML tree parser Finding elements by relationships, traversing nested content, or working with a document structure. It builds a tree and may repair malformed markup according to HTML5 parsing rules.

The external golang.org/x/net/html package supports both approaches. Its parser follows HTML5 tree-construction rules, so malformed input can lead to implied, moved, or dropped nodes rather than a tree that exactly mirrors the source text. The package assumes UTF-8 input and rejects nesting deeper than 512 elements. Check the module version selected by your project.

Parsing a document is not the same as running it in a browser. A basic HTTP client and HTML parser do not execute a page’s JavaScript. If content appears only after client-side rendering, this approach may not expose it; browser automation is a separate solution.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make link discovery safe and polite

Extracting an href is only the start of crawling. Before adding a link to a queue, resolve it against the page URL and apply your crawl policy. A small crawler should deliberately decide what it will visit rather than treating every discovered URL as permission to fetch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Limit scope: allow only the schemes and hosts you intend to crawl, and reject out-of-scope redirects or links.
  • Avoid repeats: track visited URLs and normalize or otherwise handle duplicates according to your needs.
  • Bound activity: limit concurrent work and requests per host; add cancellation so the crawl can stop cleanly.
  • Respect site rules: review the site’s terms and applicable legal requirements. RFC 9309 defines the Robots Exclusion Protocol; robots.txt coordinates crawler behavior but is not authentication, authorization, or a security boundary. See RFC 9309.

Redirects also deserve attention when requests carry credentials. Go documents that it strips the Authorization header on a redirect to a domain that is neither an exact match nor a subdomain of the original. A scraper may also need custom redirect rules to enforce its own scope; consult Go’s security decisions and the net/http documentation.

Handle common failures deliberately

  • Invalid or unsupported URL: reject input that cannot be parsed or lacks a host, and permit only schemes your scraper is designed to fetch.
  • Request error or timeout: report the error and stop or retry according to an explicit policy; do not retry indefinitely.
  • Non-success HTTP status: decide whether to skip, report, or handle the response specially. Close its body even when you do not parse it.
  • Oversized response: stop after your configured limit and treat the page as too large rather than reading the full body into memory.
  • Unexpected extraction: inspect the parsed tree and remember that HTML5 repair can change how malformed source markup is represented.
  • Missing content: check whether the data is present in the HTTP response HTML; content rendered only by JavaScript is outside this basic scraper’s capabilities.

The io package provides streaming primitives, but the application must decide how many bytes it is willing to consume. The sample’s bounded read demonstrates one practical limit.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.