What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Andrew Ng’s central lesson from 2024 was that AI progress increasingly came from how models were organized into useful workflows, not only from making foundation models larger. His year-end The Batch roundup highlighted agents, falling prices and increasingly capable smaller models, while his BUILD keynote emphasized agentic reasoning and the growing value of unstructured data.
That does not mean autonomous AI employees arrived. It means developers gained more effective patterns for combining models with tools, retrieval, planning, verification and human oversight.
Ng’s main 2024 thesis: progress moved from models to systems
Model quality still mattered in 2024. Better training, more compute and new model releases improved what applications could do. But Ng’s interpretation placed increasing emphasis on the surrounding system: the data it can access, the tools it can call, the number of steps it can take and the checks applied before an answer becomes an action.
In this view, an application can improve without waiting for a new frontier model. Retrieval can supply authoritative information, structured outputs can make responses usable by software, and a verification step can catch errors. The practical unit of progress is therefore often the complete workflow rather than a benchmark score from an isolated model.
#1 Best Overall
Ng’s year-end roundup, “Top AI Stories of 2024!”, is an interpretation of the year rather than a complete ranking of every AI event.
What “agentic workflow” means
An agentic workflow is a system in which a model performs multiple reasoning or action steps instead of producing one response and stopping. A controller supplies instructions, state and permissions; the model decides what to do next within those boundaries.
Reflection
The model produces an output, critiques it and revises it. A coding workflow might write a function, run tests, inspect a failure and submit a corrected version. Reflection can improve quality, but a confident self-critique is not proof of correctness.
Tool use
The model calls an external capability rather than relying only on information encoded in its parameters. Tools can include search, a calculator, a database, a code interpreter, browser automation or a business-software API. Typed arguments, permissions and confirmation gates are essential when a call can change data or spend money.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #2
Planning
The model creates or follows a sequence of sub-tasks. For example, a research system could gather sources, compare competitors, identify disagreements and produce a cited report. Plans should have step limits and intermediate validation so an early mistake does not contaminate every later result.
Multi-agent collaboration
Several specialized agents can divide work: one gathers evidence, another analyzes it, a third challenges the analysis and a final component formats the result. This is an architecture pattern, not evidence that the software has human-like agency. More agents also mean more calls, coordination failures and opportunities for error.
Ng discussed these patterns and the rise of agentic reasoning in his 2024 BUILD keynote, which also connected AI’s progress with text, images, video and audio: watch the keynote.
Why agents mattered in 2024
Foundation-model releases no longer determined application quality by themselves. A cheaper or less powerful model could sometimes deliver better task results when surrounded by retrieval, tools, iterative checking and a well-designed controller. This moved engineering attention toward orchestration, data quality, evaluation, latency and cost.
Free tools Windows power users keep installed
One-click scans. No signup required.
Agents are most useful when a task is variable, naturally multi-step and dependent on external information. They are usually a poor choice for deterministic work that a normal API call, SQL query or rule can perform more predictably.
| Approach | Strength | Typical risk |
|---|---|---|
| Single model call | Low latency and simple operations | Limited verification and context |
| Retrieval or tool-assisted workflow | Grounded answers and access to current systems | Bad retrieval or incorrect tool arguments |
| Planning and reflection | Better handling of complex, multi-step tasks | Higher latency, cost and possible reasoning loops |
| Multi-agent workflow | Role specialization and independent critique | Coordination overhead and cascading errors |
AI became cheaper—and smaller
Falling inference prices changed what teams could attempt. Repeated model calls, which are common in agentic workflows, became more affordable; startups could test ideas that had previously been uneconomic; and high-volume applications could consider smaller models for routine steps.
However, a lower token price is not the same as a lower workflow cost. One request may trigger several model calls, long contexts, embeddings, retrieval, browser sessions, tool execution and human review. The useful measure is cost per successful task, not merely cost per token.
Why smaller models gained importance
- They generally reduce serving cost and response time.
- They can be easier to deploy privately or on constrained hardware.
- They suit narrow, repeated tasks such as classification, extraction or routing.
- They make frequent agentic calls more economical.
Smaller models did not universally match the largest systems. Select models using task accuracy, repeated-run reliability, context needs, tool-use behavior, structured-output compliance, privacy, latency and total workflow cost. Ng’s roundup presented shrinking models as an expansion of the design space, not a declaration that large models were obsolete.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchUnstructured and multimodal data became strategic
Ng’s BUILD keynote highlighted the value of unstructured data: documents, emails, images, recordings, video and other material that traditional enterprise systems often leave difficult to search or use. Multimodal models can turn some of this material into searchable evidence, extracted fields or workflow decisions.
- Manufacturers can analyze inspection images and video.
- Support teams can search transcripts for recurring problems.
- Legal and compliance groups can review document collections.
- Researchers can work with scientific images and audio, subject to applicable rules.
- Organizations can make manuals and internal knowledge more accessible.
Multimodal is not a single reliability level. Accepting an image does not guarantee accurate counting, small-text recognition, spatial reasoning, defect detection or identity tracking across video. Unstructured data may also be stale, duplicated, biased, poorly labeled, copyrighted or restricted by privacy rules.
Reasoning, open models and competitive pressure
2024 brought greater attention to systems that spend additional computation on difficult problems. Extra reasoning steps can help, but they also add latency and cost. Teams should test whether gains hold on real tasks outside benchmarks, whether intermediate work is verifiable and whether the model still makes consequential mistakes despite sounding more deliberate.
Open and openly available models broadened experimentation and private deployment. “Open source,” “open weights” and a commercial API are different arrangements, with different levels of access to code, weights and training information. Downloadability does not automatically make a model cheaper: hosting, GPUs, security, upgrades, monitoring and engineering become the operator’s responsibility.
Best Value
Prototyping accelerated; dependable products remained hard
AI made it faster and cheaper to test ideas. A prototype can work on curated examples with manual review and tolerate occasional failure. A production system must define success, handle malformed inputs, control permissions, log actions, detect failures, protect sensitive data and provide human escalation.
Evaluate the whole workflow
- Task completion and factual correctness.
- Accuracy of tool selection and arguments.
- Number of steps, latency and cost per successful task.
- Recovery after failed calls or ambiguous requests.
- Human override and escalation rates.
- Privacy, security and policy violations.
- Reproducibility across repeated runs.
Common failure modes and controls
- Agent loops: impose maximum steps, timeouts, duplicate-action detection and budgets.
- Wrong tool calls: use strict schemas, typed arguments, permission boundaries and confirmation for consequential actions.
- Cascading errors: validate intermediate results, require evidence and support rollback.
- Prompt injection: treat retrieved pages, files and emails as untrusted data, isolate them and restrict side effects.
- Stale enterprise data: track owners and dates, rank authoritative sources and show provenance.
- False confidence: test against verified answers, display uncertainty and use deterministic checks where possible.
- Hidden cost inflation: set per-task budgets, cache reusable results and route simple steps to smaller models.
What 2024 did not prove
- Agents were generally autonomous or ready to operate without supervision.
- A larger model was always better for a particular business task.
- Benchmark gains translated directly into business value.
- Chatbots could safely perform irreversible actions on their own.
- Generated content was reliably factual.
- Agentic systems eliminated software engineering.
- Lower API prices automatically made applications profitable.
- Multimodal models understood the world like people.
- Artificial general intelligence had arrived or was demonstrably imminent.
Ng popularized and clearly articulated agentic workflow patterns; he did not invent the underlying ideas, which draw on earlier work in planning, tool use, software agents, reinforcement learning and multi-agent systems.
Practical lessons for teams in 2026
- Start with a workflow, not the word “agent.” Map the task, decisions, data and allowable actions first.
- Choose the smallest model that meets requirements. Compare quality, reliability, latency, privacy and cost on your own workload.
- Benchmark the complete task. Include retrieval, tools, prompts, retries, human review and failure recovery.
- Add tools only when they create measurable value. Every tool expands the attack surface and testing burden.
- Put approval gates around irreversible actions. Sending messages, changing records or making purchases should not rely on an unchecked model decision.
- Measure cost per successful outcome. Include all model calls, infrastructure, storage, observability and human intervention.
- Treat data quality as a product dependency. Assign ownership, track freshness, enforce access controls and expose provenance.
For structured learning, DeepLearning.AI offers Andrew Ng’s courses, while Coursera provides broader curricula and certificates. Developers can compare managed APIs from OpenAI, Anthropic and Google; Google Cloud teams may use Vertex AI. Open-model experimentation is available through Hugging Face, while LangChain and LangGraph documentation cover orchestration patterns. Industrial visual-inspection teams can investigate Landing AI. Current plans, limits and prices vary by model, geography and date.
These products solve different problems; none removes the need to define acceptable error, privacy constraints, latency, maximum cost and an escalation path.
Why this interpretation still matters
The durable lesson of Ng’s 2024 roundup is not that every application needs an autonomous agent. It is that useful AI increasingly depends on combining a suitable model with sound software engineering, trustworthy data, explicit evaluation and carefully bounded workflows. That shift—from bigger models alone to better systems—was the year’s most consequential practical change.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

