Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install GoSpider with Go, then start a crawl of a site you are authorized to assess with gospider -s "https://example.com/". Add -d to limit crawl depth, -c to control concurrent requests, and -o to save output. For a newline-delimited list of sites, use -S and set site-level parallelism with -t. GoSpider is a command-line crawler, not a browser screenshot tool; the guide below covers both its crawl workflow and the cases where a screenshot API is a better fit.

What GoSpider does—and what it does not

GoSpider is an open-source web spider written in Go. It crawls one site or a list of sites and can report discovered URLs in output designed to work with command-line tools. Its documented capabilities include JavaScript link finding, sitemap and robots.txt parsing, subdomain discovery, third-party URL sources, Burp request input, parallel crawling, and random user agents. These are discovery options, not guarantees: a crawl can only find artifacts exposed by the target and the sources you enable.

Use GoSpider when you need to discover URLs within an authorized scope. It does not return a clean PNG, JPEG, WebP, or PDF screenshot as its primary output, and crawling a URL is not equivalent to visually capturing a rendered page.

Install GoSpider and verify the binary

Install with Go

The upstream README documents installing the latest Go module version with:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
GO111MODULE=on go install github.com/jaeles-project/gospider@latest

Ensure the Go binary directory is on your shell’s PATH if the command is installed successfully but gospider is not found. Then check the CLI available to you:

gospider --help
gospider --version

Confirm the version before relying on an option from an example or another guide. The upstream README usage block displays v1.1.5, while the Kali Linux tools page displays v1.1.6; source and package versions can differ. The version reported by your installed binary is the practical reference for your setup.

Build with Docker

The project also documents a Docker workflow: clone the GoSpider repository, build an image from the repository directory, then run its help command:

docker build -t gospider:latest gospider
docker run -t gospider -h

The build command assumes the repository has been cloned into a directory named gospider and that you run the build from its parent directory. Check the project’s current instructions if your checkout layout differs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run a first crawl of one site

Start with one authorized site and a shallow crawl:

gospider -s "https://example.com/" -o output -c 10 -d 1

-s (or --site) selects one site. -o (or --output) sets an output folder. -d (or --depth) limits recursion depth; -c (or --concurrent) sets the maximum concurrent requests for matching domains. The example uses depth 1 and concurrency 10 as explicit starting settings, not universal optimal values.

For a quick test without a saved-output folder, use:

gospider -s "https://example.com/"

Review the results and confirm that discovered URLs remain within the scope you are authorized to crawl before increasing depth or broadening discovery sources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Crawl a list of sites and choose output

Use a newline-delimited site file

Put one site per line in sites.txt, then run:

gospider -S sites.txt -o output -c 10 -d 1 -t 20

-S (or --sites) reads a list of sites; -t (or --threads) controls how many sites run in parallel. This is distinct from -c, which controls concurrent requests for matching domains. Keep both values conservative when working against systems where load matters.

Pick an output mode

  • --json requests JSON output.
  • -q or --quiet suppresses other output and prints URLs, useful for piping into another command.
  • -v or --verbose emits verbose logs.
  • -l or --length shows response length.
  • -L or --filter-length filters by response lengths.
  • -R or --raw selects raw output.

Check gospider --help for the exact behavior and accepted values in your installed version, especially when combining output options.

Set depth, concurrency, timeout, and delay responsibly

For a first pass, use a shallow depth such as -d 1, modest concurrency such as -c 5 or -c 10, and a timeout suited to your network and target. The README documents a default concurrency of 5 and a default request timeout of 10 seconds; these are program defaults, not speed or capacity recommendations. -m or --timeout sets request timeout in seconds.

The README says depth 0 means infinite recursion. Treat that setting as potentially expansive: use it only when authorization and scope clearly permit an unbounded crawl, and when you have a plan to stop or constrain the run. --delay adds a fixed delay between requests; --random-delay adds randomized delay. Add delay or reduce concurrency when you need to limit request rate. The source material does not establish a universal crawl speed, so tune against your environment rather than expecting a fixed completion time.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The README examples include --blacklist for URL regular expressions and note that common static file extensions are filtered by default. Use a blacklist to exclude known out-of-scope URL patterns, and verify the pattern against representative URLs before a long run.

Supply headers, cookies, a proxy, or a Burp request

For an authorized authenticated crawl, GoSpider accepts repeated headers with -H and cookies with --cookie. For example:

gospider -s "https://example.com/" -H "Accept: */*" -H "Test: test" --cookie "testA=a; testB=b"

The example values are illustrative; substitute only credentials and headers approved for the engagement. The CLI also documents:

  • -p or --proxy to use a proxy.
  • -u or --user-agent for built-in random web/mobile agents or a custom user-agent string.
  • --burp burp_req.txt to load headers and cookies from a raw Burp request.

Keep request files and credentials out of shared logs and repositories. Restrict targets to the approved scope; proxy use does not change what you are authorized to access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Expand URL discovery selectively

Enable only the discovery sources that answer your question. The README lists these controls:

  • --js enables JavaScript link finding.
  • --sitemap tries sitemap.xml.
  • --robots tries robots.txt.
  • --subs includes subdomains.
  • --other-source obtains URLs from Archive.org, Common Crawl, VirusTotal, and AlienVault.
  • --include-subs and --include-other-source broaden how those discovered URLs are incorporated.

The project also lists AWS S3 references and link-finder behavior among its features. Enabling JavaScript, subdomains, or third-party sources can substantially broaden what appears in results; inspect the URLs and keep the actual crawl within your authorization. A source reporting a URL does not establish that the destination is in scope or currently accessible.

Troubleshoot common problems

gospider is not found

The installation may have succeeded while the Go binary directory is absent from PATH. Check the Go installation environment, add its binary directory to PATH, reopen the shell if needed, then run gospider --version.

The command rejects an option

Option availability can depend on the binary or package version. Run gospider --help on the installed executable and use its displayed flags rather than assuming every version matches an online example.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The crawl returns few or no URLs

Check that the site is reachable from the machine running GoSpider, that the start URL is correct, and that the chosen depth is not too restrictive. Try the target’s sitemap or JavaScript discovery only when relevant, and inspect verbose output with -v to understand what the run reports. GoSpider’s documented features cannot guarantee that a target exposes discoverable links.

Requests time out or the target responds poorly

Increase -m only when longer response waits make sense; otherwise lower -c and add --delay or --random-delay. A longer timeout can also lengthen a run when endpoints do not respond, so it is not a general fix for a crawl that stalls.

Authenticated pages are missing

Verify that the supplied cookie or headers are valid for the target and that the Burp request file contains the request data you intend to reuse. Do not paste secrets into public logs or share request files beyond the authorized team.

Results include out-of-scope URLs

Discovery sources, subdomains, and links can surface destinations beyond the starting host. Disable broad sources you do not need, add a URL-regex blacklist where appropriate, and review outputs before using them for further requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

GoSpider discovers URLs; it is not a screenshot API. If your actual task is to capture a page as an image or PDF rather than enumerate links, ScreenshotNeo is a separate option: one GET request can return a screenshot or PDF without setting up browser automation locally. For example, this cURL request captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, and cache hits are not billed. It also provides an MCP server for AI agents, and its free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. These are screenshot-service capabilities, not GoSpider crawling features. Sign up free for 1,000 screenshots a month with no card.

Choose settings by the job

Need Start with Why
Explore one authorized site -s, shallow -d, modest -c Limits the initial scope and request pressure.
Process several approved sites -S with -t; tune -c separately Site-level parallelism and per-domain request concurrency are different controls.
Inspect a logged-in area -H, --cookie, or --burp Supplies authorized request context to the crawl.
Find URLs beyond ordinary links Enable only needed options such as --js, --sitemap, or --other-source Broad discovery may surface URLs outside the starting site or scope.
Capture a rendered page as an image or PDF Use a screenshot tool rather than a crawler A URL list and a visual capture are different outputs.

Frequently Asked Questions

Does GoSpider crawl an entire website automatically?

Not necessarily. Crawl coverage depends on its configured depth, available links, enabled discovery sources, and what the target exposes.

What does GoSpider’s depth value of 0 mean?

The upstream README describes 0 as infinite recursion; use it only when an unbounded crawl is appropriate for the authorized scope.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can GoSpider take a screenshot of each page?

Its documented purpose is URL crawling and discovery, not returning page screenshots. Use a screenshot capture service when the required output is an image or PDF.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.