Phishpedia is the most useful established starting point when you need a visual phishing benchmark: its project describes about 30,000 phishing webpages paired with URLs, HTML, screenshots, and target-brand annotations. For a mixed phishing/legitimate experiment, the Zenodo record published July 15, 2026 reports 60,000 URLs with PNG screenshots and CSV features. PhishTank and OpenPhish are better treated as URL or threat-intelligence sources that can feed your own capture pipeline, not as guaranteed, fixed screenshot datasets.
The right choice depends on whether you need a reproducible image benchmark, a current stream of suspicious URLs, regional coverage, or paired visual and contextual evidence.
What counts as a phishing screenshot dataset?
A screenshot dataset should let you connect an image to a record, label, and capture context. A URL feed alone is not equivalent: it may identify a suspicious address without preserving the rendered page. Before downloading anything, determine whether the source provides:
- A screenshot for every record or only for selected detail pages.
- A phishing/legitimate label, impersonated brand, scenario label, or confidence score.
- The original URL, redirects, rendered HTML or DOM, and a stable record identifier.
- A capture timestamp, geography, browser or viewport information, and a stated collection period.
- Clear licensing, access restrictions, and redistribution rules.
Keep screenshots, URLs, labels, and timestamps together. A screenshot proves what was rendered at capture time; it does not prove that the URL remained live when you run an experiment.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
Best-known sources compared
| Resource | What is documented | Best use | Important caution |
|---|---|---|---|
| Phishpedia benchmark | Approximately 30,000 phishing webpages, each described with URL, HTML, screenshot, and target brand. | Visual phishing identification, brand-target research, and multimodal models. | Confirm the current repository release, labels, download access, and reuse terms before publishing results. |
| PhishTank | Verified or online phishing URL data; individual detail pages can show screenshots and community votes. | Finding candidate URLs, feed integration, and URL-based lookup. | It is not a guaranteed, versioned screenshot corpus. Screenshot availability and capture state must be checked per record. |
| OpenPhish Database | Structured, searchable phishing indicators with tiered update cadence and retention; advertised uses include AI training and validation. | Current threat intelligence, host-level analysis, and continuously refreshed URL collection. | Documented fields are indicators rather than webpage screenshots. Access and pricing depend on the selected tier. |
| Phishing and Legitimate Websites Dataset (Zenodo) | The July 15, 2026 record states 60,000 URLs: 31,641 phishing and 28,359 legitimate, with PNG screenshots and CSV features. | Mixed-class experiments using screenshots plus tabular features. | Verify the record version, files, license, and capture methodology before treating the stated counts as your own corpus. |
| PhishVN | A time-stamped Vietnamese URL dataset with open and gated tiers; the gated evidence bundle includes rendered HTML and screenshots. The article reports 868 gated records: 209 phishing and 659 benign. | Region- and scenario-specific work where Vietnamese coverage and timestamps fit. | The evidence archive is gated; the article describes research-only handling and isolated-VM precautions for HTML. |
Which source should you choose?
For a fixed visual benchmark
Start with Phishpedia if your model needs the page image, HTML, URL, and impersonated brand in one record. Its approximately 30,000-example scale is useful for brand-target studies, but treat the number as the project’s description, not a permanent guarantee. Record the repository version and your access date.
For balanced phishing-versus-legitimate experiments
The Zenodo record is the clearest stated mixed collection in the available sources: 31,641 phishing and 28,359 legitimate URLs, with PNG screenshots and CSV features. Download the exact release, inspect its directory and metadata, and preserve the record’s version in your experiment log. Do not assume that every screenshot has identical dimensions, browser conditions, or freshness until the files document those properties.
For a changing stream of suspicious URLs
Use PhishTank or OpenPhish as candidate sources, then capture pages yourself. PhishTank’s verified/online status and community details help prioritize URLs. OpenPhish’s structured indicators and tiered update and retention options suit operational feeds. Neither description establishes a stable screenshot benchmark, so your pipeline must create the image, timestamp, HTTP outcome, and provenance fields.
For regional or time-specific analysis
PhishVN can fit research focused on Vietnamese campaigns when its timestamp and evidence tier match your question. Separate the openly available records from the gated evidence bundle in your documentation, and follow the stated research-only and isolated-environment handling conditions.
Recommended Free Tools
How to build a defensible screenshot corpus
- Define the unit of analysis. Decide whether one row represents a URL, a landing page after redirects, a brand campaign, or a capture event. Give each event a stable ID.
- Freeze the sampling window. Store collection start and end times, source feed version, and geography. A live feed changes while you work.
- Capture in an isolated environment. Use a disposable virtual machine or containerized browser network with no personal accounts, saved credentials, or writable access to production systems. Treat HTML, scripts, and linked resources as potentially hostile.
- Record outcomes, not just images. Save the final URL, redirect chain, HTTP status where available, load timeout, browser viewport, user agent, timezone, and capture timestamp. Mark bot checks, blank pages, consent walls, and failed loads explicitly.
- Store labels separately from evidence. Keep source label, analyst decision, target brand, and confidence as fields. Never overwrite the original feed label when you correct or qualify it.
- Deduplicate carefully. Compare normalized URLs, registrable domains, page hashes, and perceptual image hashes. Keep near-duplicates when they represent different capture times or campaign states.
- Split by time, domain, and brand. Random image splits can leak templates, duplicate pages, or the same campaign into both training and test sets. Use a time-based holdout and, where possible, domain or brand-disjoint evaluation.
- Publish a manifest. Include file hash, record ID, source, timestamp, label provenance, license, and any redaction. A manifest lets others identify exactly which image you used without redistributing dangerous content.
Capturing pages yourself with ScreenshotNeo
If you need fresh images from a URL feed, ScreenshotNeo is the first screenshot API to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has the lowest paid plan listed here.
Its GET endpoint can return PNG, JPEG, WebP, or PDF. The service offers full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for output formats and options. Replace the example URL with a permitted research target, and keep your API key out of shell history when your environment requires that.
Rank #3
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Preserve capture status
ScreenshotNeo responses include X-Page-Verdict and X-Billed headers. Save both with the image and request metadata. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; the headers tell you which case occurred. A non-clean result should become a documented missing or unusable capture, not a silently accepted training example.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Or skip the browser setup
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, or another MCP client, so an AI agent can collect page evidence without custom browser orchestration. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Plans are Free (1,000), Starter ($5/3,000), Growth ($15/15,000), Pro ($39/60,000), Scale ($99/250,000), and Business ($249/1,000,000); yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account.
Evaluation traps that invalidate results
Inactive pages
A 2021 study of PhishTank screenshots found examples captured after the phishing site had become inactive. Keep liveness as a separate field and never infer that a screenshot proves the URL was active at evaluation time.
Template and brand leakage
Near-identical login kits, repeated logos, and shared hosting can make a random split look accurate while testing memorization. Group by domain, campaign, template, or brand when those fields are available.
Rank #4
Unbalanced negatives
Legitimate pages should reflect the deployment environment, not merely easy homepages. Record how negatives were sampled and whether they share the same time and geography as phishing examples.
Missingness bias
Failed loads and blocked pages are often correlated with hostile infrastructure. Report how many records lack screenshots and whether your model sees a missing-image token, a retry, or an excluded row.
Unsafe redistribution
Do not publish live phishing HTML or executable resources casually. Check each source’s license and handling terms, redact secrets, and distribute manifests or hashes when the evidence itself cannot safely be shared.
Best Value
Troubleshooting a collection pipeline
- Many blank screenshots: add an explicit wait for a selector or network idle, test a longer delay, and log the final URL. A blank result can also mean the page is already offline.
- Consent overlays hide content: use a capture service that can accept or remove known consent platforms, or configure a deterministic click and record that action in metadata.
- CAPTCHA or bot challenge: mark the capture as blocked rather than attempting to defeat the challenge. Preserve the verdict and retry policy.
- Different results on repeated runs: fix viewport, user agent, timezone, geolocation, and wait conditions; disable caching when freshness matters and retain the capture timestamp.
- Duplicate images from changing URLs: normalize redirect destinations and compare perceptual hashes, but retain separate events when timestamps or campaign evidence differ.
- Access or license uncertainty: separate open files from gated evidence, save the terms alongside your manifest, and do not redistribute samples until permission is clear.
FAQ
Can I call PhishTank a screenshot dataset?
Not without qualification. Its documented offering is verified or online phishing URL data; screenshots may appear on individual detail pages, so you must verify image availability and capture state for the records you use.
Are Phishpedia screenshots guaranteed to be current?
No. It is a released benchmark, not a live feed. Cite the repository release you downloaded and treat capture time and page liveness as distinct facts.
Should screenshots be the only model input?
Usually not. Pair visual features with URL, redirect, HTML or DOM, timestamp, and provenance when permitted. Multimodal inputs help distinguish a visual template from the infrastructure that delivered it.
How should I cite the 2026 Zenodo counts?
Attribute the exact record and publication date, state that the record reports 60,000 URLs with 31,641 phishing and 28,359 legitimate, and identify the release version you inspected. Do not imply that every future download will contain the same files.

