Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with the decision your model must make, then collect examples that represent the people, conditions, and edge cases it will encounter. Define what each example contains, how its answer will be labeled, and whether you have permission to use the data. There is no universal number of rows that makes a dataset sufficient: coverage, label quality, evaluation results, and ongoing monitoring matter more than raw volume.

1. Define what the model needs to learn

Before gathering data, write down the intended use of the model and the decision it should support. Translate that into a data specification: what one example represents, what information is available when the model makes its decision, and what outcome it should predict. AWS describes a supervised-learning example as having a target and variables, or features, that help predict it.

Specify the target and unit of observation

The target, also called the label, is the answer the model is meant to predict. Features are the observed attributes it can use to make that prediction. For a model that flags a potentially fraudulent transaction, one row might represent a transaction; the target could be whether it was later confirmed as fraud. The features must be information legitimately available at the moment the system would assess that transaction.

Define the unit consistently. Mixing customer-level and transaction-level records without an explicit design can create ambiguous labels and misleading counts. Also define how the target is established, the time window for establishing it, and what to do when the outcome is unknown or disputed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Set the operating conditions and error trade-offs

List the users, locations, languages, devices, time periods, and other conditions in which the system is expected to operate. Identify errors that are especially costly: a false alarm, a missed event, a delayed result, or an incorrect recommendation may have different consequences. Those requirements determine what examples and labels you need, and how you should evaluate the finished model.

Write down the intended use and material exclusions. A dataset suitable for an internal research prototype may not be suitable for a consequential decision about people. Narrowing the task can make data collection more achievable and reduce the risk of using information for a purpose it was not collected to support.

2. Choose a collection approach and source

There is no single best source. Compare what each source can provide against coverage, label availability, provenance, permissions, privacy and security risks, freshness, expected label error, and total operational cost. OECD’s 2025 report, Mapping relevant data collection mechanisms for AI training, emphasizes that collection mechanisms have different implications for developers, data subjects, and other rights holders.

Approach Useful when Check before relying on it
Existing labeled dataset You need a starting point or a baseline and the dataset matches the task closely. Confirm its collection purpose, permissions, population, label definitions, dates, transformations, and known gaps. A dataset being available does not establish that it is appropriate for your use.
Operational records Your organization already records events related to the target. Check whether the records reflect actual outcomes or merely past decisions. Missing records, policy changes, or a process that treated groups differently can shape the labels.
Direct contributions You need people to provide examples, answers, ratings, or other information. Explain the purpose and collection conditions; consider who is likely to volunteer and who is not. Plan for consent, contributor treatment, and quality review.
Observed or acquired data Relevant behavior or content can be observed or obtained from a supplier. Document the source, collection method, geographic and population coverage, rights, restrictions, and permitted purpose. Assess supplier and downstream risks.
Newly collected data Existing sources do not provide needed coverage, labels, or conditions. Plan the sampling, collection environment, labeling workflow, protections, cost, and update cadence before collection begins.

Google’s People + AI Guidebook recommends weighing predictive power, relevance, fairness, privacy, and security when deciding whether to use an existing dataset or create one. If you use web pages as visual input—for example, to train a model that recognizes page layouts—first confirm that the pages may be captured and used for the intended purpose. A screenshot records page content that may include personal or sensitive information; treat it as data requiring governance, not as a permission-free image.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Document provenance, purpose, and permissions

Keep a record with the dataset, not just in a project conversation. For each source, capture:

  • Who collected or supplied it, and when and where collection occurred.
  • The collection method, population or coverage, and the original purpose.
  • The permissions, restrictions, and applicable legal basis for the intended use.
  • Transformations, filtering, labeling, and joins performed after collection.
  • Known omissions, quality issues, geographic limits, and changes over time.
  • Who can access the data, for what purpose, and how access decisions are logged.

For personal data, have an appropriate legal and privacy review establish the lawful basis and communicate the purpose to people as required. Microsoft’s Azure Machine Learning guidance says, “Obtain voluntary informed consent.” It also advises using data only for purposes covered by the documented consent, retaining consent records, qualifying suppliers and geographies, and stewarding datasets. Consent is not a blanket permission for any later use; confirm the conditions that apply to the particular data and jurisdiction.

Minimize the information collected to what the task requires. Use access controls and encryption, and consider de-identification or pseudonymisation where appropriate; these measures reduce exposure but do not automatically eliminate re-identification risk or other obligations. The UK National Cyber Security Centre identifies filtering, sanitisation, differential privacy, masking, aggregation, swapping, and pseudonymisation as possible controls. Choose controls based on the data, threat model, and use rather than treating any one technique as a universal solution.

4. Collect examples that reflect actual use

Sample across the operating range you defined, including relevant edge cases and subgroups. AWS guidance stresses that examples should represent cases the model will encounter, including positive and negative cases where relevant. Google’s guide likewise emphasizes relevance and fairness. A large convenient sample cannot repair a systematic omission: if a language, device, region, or group is absent, adding more examples from already-covered conditions will not fill that gap.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make coverage visible

Before collection, map expected cases to available cases. Track useful dimensions such as time period, location, language, device, event type, or subgroup, when those dimensions are appropriate and can be handled lawfully. Look for combinations as well as individual categories: a dataset may include examples from two groups separately but very few examples where their conditions overlap.

Record collection context, since changes in a form, sensor, policy, user interface, or operating environment can change what a feature means. If important cases are rare, use a planned strategy to obtain or review more of them; do not let an apparently balanced total conceal gaps in the specific cases where errors matter.

Separate collection from assumptions

Observed outcomes are not always objective ground truth. An operational record may reflect a human decision, historical policy, or whether someone had the opportunity to report an outcome. Document how a label came to exist and whether that process could systematically miss or misclassify cases. Where the truth is uncertain, preserve uncertainty rather than forcing every example into a confident class.

5. Design labels and measure their quality

For supervised learning, a label is the target answer, not another feature. Google’s People + AI Guidebook notes that accurate labels are crucial and that both labeler instructions and interface design affect quality. Treat the labeling scheme as part of the data design.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write instructions before scaling annotation

Define every class or target in plain language, include positive and negative examples, and explain borderline cases, missing information, and escalation rules. Clarify whether labelers should infer, transcribe, judge, or report only what is directly observable. Pilot the instructions on a small varied sample and revise them when labelers repeatedly disagree for the same reason.

Review disagreement and errors

Train labelers on the final instructions and provide a way to flag ambiguous examples. Measure agreement or error using an appropriate review process; investigate disagreement rather than assuming that majority vote always reveals the truth. Use expert review for cases requiring expertise, and preserve adjudication decisions so the team can understand how final labels were reached.

Where you hire outside help, compare data-labeling services, human data collection platforms, or managed annotation against internal work on instruction control, reviewer qualifications, security, privacy, provenance, contributor conditions, expected error, and total cost. Do not outsource responsibility for whether the labels are fit for the model’s intended use.

6. Run data-quality checks before training

Quality is multidimensional and should be checked before training and as data changes. The UK Data and AI Ethics Framework names completeness, accuracy, validity, consistency, uniqueness, and timeliness among the relevant dimensions. Add checks for duplicates, missingness, outliers, class balance, leakage, and subgroup coverage where they apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Completeness: Are required fields or cases missing, and is missingness concentrated in particular groups or conditions?
  • Accuracy and validity: Do values and labels match the definitions and permissible ranges for the task?
  • Consistency and uniqueness: Are equivalent records represented consistently, and are duplicate examples inflating apparent coverage?
  • Timeliness: Does the data still reflect the current system and deployment conditions?
  • Balance and coverage: Are important outcomes, subgroups, and edge cases represented well enough to evaluate?
  • Leakage: Does any feature reveal the target or information that would not be available at decision time?

Make checks repeatable rather than relying on a one-time visual inspection. Keep counts and summaries by relevant subgroup and time period, while applying appropriate privacy controls. A check should produce a decision: accept, investigate, correct, exclude with a documented reason, or collect more data.

Example: a small Python CSV check

This standard-library script checks that a CSV has the expected columns, contains non-empty IDs and labels, and has no duplicate IDs. Change the column names to match the dataset. It is a first-pass integrity check, not a substitute for reviewing label validity, coverage, leakage, or legal permissions.

import csv
from collections import Counter

path = "training_data.csv"
required = {"example_id", "label", "feature_text"}
ids = []
rows = 0
missing = Counter()

with open(path, newline="", encoding="utf-8") as f:
    reader = csv.DictReader(f)
    headers = set(reader.fieldnames or [])
    absent = required - headers
    if absent:
        raise SystemExit(f"Missing required columns: {sorted(absent)}")

    for line_number, row in enumerate(reader, start=2):
        rows += 1
        for column in required:
            if not (row.get(column) or "").strip():
                missing[column] += 1
        ids.append((row.get("example_id") or "").strip())

duplicates = [value for value, count in Counter(ids).items()
              if value and count > 1]
print(f"Rows: {rows}")
print(f"Blank required values: {dict(missing)}")
print(f"Duplicate nonblank IDs: {len(duplicates)}")
if not rows or missing or duplicates:
    raise SystemExit("Review data-quality issues before training")

For a real pipeline, add task-specific validation, such as permitted label values, date ranges, duplicate detection across content, or coverage summaries. Avoid printing personal data into logs while diagnosing failures.

7. Split, version, and preserve lineage

Keep training, validation, and test data separate according to the evaluation design. The test set should represent the intended evaluation conditions; it should not influence model fitting or iterative choices if you want it to provide an independent final assessment. Prevent duplicates and future information from leaking across splits. For time-dependent use, a time-aware split may better reflect predicting future cases than a random split.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Version raw and transformed datasets, labels, instructions, and code or rules used to prepare them. Preserve lineage from source through each transformation, including exclusions and label changes. The UK AI-ready dataset guidance recommends metadata, stewardship, transformation documentation, catalogs, access controls, audit logging, and continuous quality monitoring. These records let a team explain which data produced a model and reproduce a dataset version when a result needs investigation.

8. Decide whether you have enough data

No authoritative numeric threshold applies to every machine-learning task. The amount needed depends on the problem, variability of cases, rarity of outcomes, label reliability, model, and intended performance. Do not treat a minimum row count found for another task as a guarantee.

Instead, justify adequacy with evidence: the required operating conditions and subgroups are covered; labels have been reviewed; the evaluation set is separated and representative of intended use; performance is acceptable for the defined error trade-offs; and uncertainty is understood for sparse cases. If results are unstable or subgroup coverage is weak, collect targeted examples or narrow the intended use rather than assuming that multiplying the dataset uniformly will help.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Monitor data after release

Collection does not end when the first model ships. Track missingness, changes in labels or collection processes, distribution shift, subgroup performance, and data drift. Set a review path for unexpected changes: identify whether the source, population, product, policy, or labeling process changed before deciding to retrain.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use governed updates. Record the new data’s provenance and permissions, rerun quality checks, preserve dataset versions, and evaluate the model against an appropriate holdout before deployment. Monitoring should also detect whether the system’s real-world use has moved beyond the purpose or population the dataset was designed to support.

10. Collect webpage screenshots for visual ML data

If your task genuinely needs webpage appearance—such as learning to classify layouts or identify visual components—a screenshot can be one input type. Decide which pages and states are needed, how to handle dynamic content and consent notices, and whether collection and model use are permitted. A screenshot may contain personal information or third-party content; review the source’s terms and applicable requirements, minimize capture to what is necessary, and apply appropriate retention and access controls.

Capture one page with ScreenshotNeo

For webpage image collection, ScreenshotNeo provides a screenshot API. The example below captures a target page as WebP; replace the URL with a page you are authorized to collect. Keep your API key private. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request can be made from Python or Node.js:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For this use case, plan consistent capture conditions: viewport or device, wait behavior, full-page versus viewport capture, and output format. A change in capture settings can create apparent variation unrelated to the page itself, so record the settings with the dataset version. ScreenshotNeo accepts parameters for full-page capture, element selection, viewport and device presets, dark mode, custom CSS or JavaScript, waiting for a selector or network idle, and other capture controls. Do not assume a capture should bypass access controls or permission requirements.

Or skip the browser setup

ScreenshotNeo can capture a page with one GET request as shown above. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server includes tools for AI agents to take screenshots. Free includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. These are capture-service features, not a replacement for dataset permissions, sampling, labeling, or quality checks. Learn about ScreenshotNeo, or sign up free for 1,000 screenshots a month with no card.

Frequently Asked Questions

Can unlabeled data be used for machine learning?

Yes. It can be useful for some approaches, but it does not provide the target answers required for ordinary supervised training. Whether it helps depends on the learning method and the task.

Should I keep examples that labelers disagree about?

Do not silently force uncertain cases into a class. Flag them, investigate the source of disagreement, and decide through documented rules whether to adjudicate, preserve uncertainty, or exclude them.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.