Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Machine learning (ML) is a way of building software that learns patterns from data and uses those patterns to predict a number, assign a category, group similar cases, or generate new content. A model’s output is only useful in relation to three things: the question it was built to answer, the data it learned from, and the decision a person or organization will make with it.

What machine learning means

The U.S. National Institute of Standards and Technology (NIST) defines machine learning in its glossary as “The development and use of computer systems that adapt and learn from data with the goal of improving accuracy” (NIST CSRC glossary entry for machine learning, which cites NIST SP 800-55v1). The key idea is that the system’s behavior is shaped by examples rather than written out rule by rule.

Google for Developers describes ML more operationally: training software, called a model, to make predictions or generate content from data (Google for Developers, “What is Machine Learning?”). In that guide’s words, “ML powers some of the most important technologies we use, from translation apps to autonomous vehicles.” The sentence is the institutional voice of Google for Developers; the page does not name an individual author.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the work moves from problem to decision

A machine learning project follows a path that is easy to state and easy to get wrong. Each stage shapes the next, so a weak step early on limits everything downstream.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  1. Frame the problem. Decide what the model should output and why. “Estimate tomorrow’s rainfall in millimeters” and “flag messages that are probably spam” are different tasks, and they need different data and different measures of success.
  2. Gather and prepare data. Collect examples, clean them, and decide which measurements (features) the model will see. NIST’s technical framework lists preprocessing and feature engineering as core parts of model development (NIST SP 1321).
  3. Choose and train a model. Select a method, then let it adjust its internal parameters so that its outputs match the examples as well as possible. Tuning settings such as model complexity is part of this step.
  4. Produce an output. Given new input, the trained model returns a prediction, a category, a group assignment, or generated content.
  5. Evaluate on data it has not seen. Test performance on examples held back from training, and check whether the measure used actually reflects the real-world goal. A model can score well on a convenient metric and still fail the people it is meant to serve.
  6. Put the output to human use. Decide who reads the output, what they are allowed to do with it, and who is accountable if it is wrong.

A simple illustration of steps 2 through 4 is rainfall prediction. Past weather observations serve as input data. Training lets the model learn relationships between observed conditions and the rainfall that followed. Current weather readings then become the input for a numeric forecast. The chain is illustrative: real forecasting systems involve far more data and expert checks than this summary shows.

The main kinds of machine learning

Google for Developers separates ML into supervised learning, unsupervised learning, reinforcement learning, and generative AI. The distinction that matters most in practice is what kind of output the task needs and whether the training examples come with known answers.

Approach What it learns from What the output looks like Examples named in the source
Supervised: regression Labeled examples with known answers A numeric value Estimates of house prices and travel times; rainfall prediction
Supervised: classification Labeled examples with known categories A category label Spam detection; image categorization
Unsupervised: clustering Unlabeled data Groups of similar records Finding groupings in data with no predefined answers
Reinforcement learning Feedback from actions taken in an environment Choices that improve over repeated trials Described generally; no specific application named in the source
Generative models Patterns in existing content New text, images, audio, or video Text completion, article summaries, translation, generated images

Two cautions follow from the table. First, clustering finds groups, but the groups do not explain their own meaning; a person still has to interpret them. Second, a generated output is a learned pattern applied to a new request, not a verified statement of fact.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Everyday examples, read as task types

The examples below are useful for recognizing the task type behind a familiar product. They are not evidence that any particular product performs well in every setting.

  • Numeric forecasts. Rainfall and travel-time estimates are regression tasks: the system predicts a quantity.
  • Spam filtering. Deciding whether a message is spam is a classification task.
  • Recommendations. Song or product suggestions personalize what a user sees. Google’s courses treat recommendation systems as a distinct topic in their own right.
  • Translation and text completion. These rely on models that learn language patterns to produce new text.
  • Article summaries and generated images. These are generative tasks: the model produces content that did not exist in its training set in that exact form.

NIST SP 1321 offers a more specialized set of examples from structural engineering and natural hazards. These include structural-response prediction, surrogate modeling, design optimization, hazard forecasting, structural-health monitoring, predictive maintenance, classification of disaster-reconnaissance data, and fragility-model development. The same document notes that data availability and privacy concerns have affected adoption in these fields. Listing these uses describes where ML is being explored and applied; it does not show that these problems have been solved.

Using ML output to make decisions

A prediction is not the same thing as a decision. A spam score, a risk estimate, or a recommended action becomes a decision only when a person, a rule, or an automated process acts on it. Before a model informs a choice, it helps to ask the same questions about any two candidate approaches:

  • Output and task. Does the decision need a number, a category, a group, a chosen action, or new content?
  • Data needs. Are labeled examples available, or only unlabeled ones? Is there enough data, and does it cover the situations the model will meet?
  • Evaluation. Was performance measured on data the model did not train on, and does the metric match what the organization actually cares about?
  • Interpretability and accountability. Can people understand why the model produced an output, and is it clear who answers for decisions made with it?
  • Operational fit. Does the use case raise data privacy issues, require significant computing resources, or fit into the way people already work?

The interpretability question deserves particular attention. NIST’s framework notes that transparency matters most where interpretability and accountability are paramount, and that explainability methods may not fully make complex models interpretable. It contrasts complex models with simpler decision trees, which are transparent by construction. A decision tree may be the better choice for decision support even when a complex model scores higher on accuracy, because people can follow and challenge its reasoning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What data-driven does not guarantee

“Data-driven” is a description of method, not a guarantee of quality. A model learns whatever patterns are present in its data, including gaps, errors, and historical imbalances. NIST discusses data quality and bias avoidance as parts of model development, and that is the right starting point for reviewing any system that affects people. Learning from data does not by itself ensure accuracy, objectivity, causation, fairness, or privacy. Those properties have to be checked in the specific context where the model is used.

Readers evaluating a claim about an ML system should ask a few direct questions:

  • What exactly does the model output, and what decision does that output feed?
  • What data was it trained on, and does that data match the setting where it will be used?
  • How was performance measured, and on what held-back data?
  • What happens when the model is wrong, and who is responsible for catching errors?

Where to go next

For a conceptual foundation, Google for Developers maintains introductory and advanced ML courses and guides covering problem framing, project management, clustering, recommendation systems, and responsible AI (Google for Developers, Machine Learning course catalog). Readers who want hands-on technical practice may find Jason Bell’s Machine Learning: Hands-On for Developers and Technical Professionals (second edition, John Wiley & Sons, 2020, 432 pages, ISBN 9781119642145) useful; its catalog record describes examples across ML variants, data preparation, algorithms, text, images, and streaming systems. It is written for developers and technical professionals, so it is not necessarily the best first step for a general reader. Bibliographic details are available on the Google Books record.

The most useful habit is to keep the chain in view: a clear question, suitable data, a model matched to the task, an output tested on unseen cases, and a human who understands what the output can and cannot support.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.