Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

No. An AI agent that can work on a task for longer has demonstrated a particular kind of capability; that alone does not show it learns from experience or is generally intelligent. The distinction matters because the organizations that control how models are updated also shape what those systems learn. Dr. Yichuan Zhang, CEO of Boltzbit, makes that argument in a September 30, 2026 essay for The AI Journal. His proposed alternative—learning from context at the point of use—is a company position, not an independently established solution.

What does an AI task horizon measure?

METR’s task-completion time horizon describes how long a task takes a human expert and the probability that an AI model or agent will complete it at a chosen reliability threshold. It is a benchmark of performance on tasks of a given duration—not a direct measurement of general intelligence, learning, or understanding.

That distinction is important when interpreting claims that AI can handle increasingly long tasks. A longer horizon means an agent is more likely to complete certain extended tasks under the benchmark’s conditions. It does not establish that the agent will retain what it learned, transfer skills to a different domain, or improve through repeated use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The benchmark has a defined scope

METR’s current task suite contains more than 100 software tasks and focuses primarily on software engineering, machine learning, and cybersecurity. METR cautions that performance may differ across domains and that, with the current suite, measurements above 16 hours are unreliable. These limits make task horizon useful for tracking a specific capability, but not a universal intelligence score. METR’s task-horizon methodology

Why longer task performance is not the same as learning

In his essay, Zhang summarizes the distinction as: “Autonomy is not the same as learning.” An agent might keep working for longer by planning, calling tools, or retrying actions without acquiring durable skills from those experiences. Conversely, a system could learn something useful without being able to sustain a long sequence of autonomous actions.

This is a conceptual distinction, not a result proved by the task-horizon benchmark. Horizon measures whether an agent completes tasks at a given length and reliability; it does not by itself test whether the agent’s underlying capabilities changed through use. Zhang’s further claim that agents can lose the original goal or enter loops is part of his essay’s argument; the material cited here does not establish a general rate for those failures.

What the task-horizon trend does—and does not—show

METR’s 2025 analysis estimated that the long-run frontier task-horizon trend had doubled about every seven months. A separate cross-domain analysis said the estimated interval may have shortened to about four months during 2024. These are historical estimates from benchmark analysis, not a forecast that AI will reach AGI on a schedule. METR’s 2025 task-horizon analysis and cross-domain analysis

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The trend should be read alongside the benchmark’s scope: the tasks are concentrated in a few technical fields, and the suite becomes unreliable at very long durations. A rising horizon is evidence of progress on measured tasks. It does not answer whether systems can learn from deployment, generalize broadly, or exercise judgment across unfamiliar situations.

Who controls how AI models evolve?

Zhang’s governance concern is that influence over AI’s evolution may be concentrated among the organizations able to build and retrain large models. In his account, deployed models generally do not change what they know until their owners update or retrain them. The cost and control implications are his characterization of current model-development economics, not a universal fact established by the benchmark sources.

The concern remains worth separating from any particular architecture: when a model changes, who decides what data or feedback informs the change, who can inspect it, and who is accountable for its effects? These questions become especially consequential when AI is used in areas where errors affect people, services, or institutions.

Central retraining and point-of-use learning are different proposals

Zhang contrasts centrally retrained models with “context-centric intelligence,” which he advocates as learning at the point of use. He names Boltzbit’s General Learning Intelligence (GLI) as an example. Boltzbit describes GLI as user-owned, trainable, and controllable; those are company descriptions and claims, not independent validation that the approach works as proposed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Question Central retraining, as characterized in the essay Point-of-use learning, as proposed by Zhang and Boltzbit
Where do updates happen? At central retraining, followed by distribution of an updated model. At deployment, using local organizational context or interaction.
Who is intended to direct updates? The model provider or central model owner, in Zhang’s account. The deploying organization or user, in Boltzbit’s stated positioning.
What evidence is available here? The essay’s general description of the architecture; not independently verified as a universal account. Company descriptions and research claims; not independently validated in the sources cited here.
What should be examined? Update cadence, cost, data access, and auditability. What is learned, what data is retained, how updates are evaluated or reversed, and who is accountable.

This is a comparison of design ideas, not proof that all systems in either category behave alike. A locally updated system may shift control closer to its users, but that does not automatically settle questions about data governance, reliability, or oversight.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What adoption figures can—and cannot—tell us

Wider use of AI makes the governance question more immediate, but adoption is not evidence that systems are learning. McKinsey’s 2025 survey chart says the share of respondents reporting AI use in at least one business function was 88% in 2025, 72% in 2024, and 55% in 2023. McKinsey notes that its definition of organizational AI use evolved over time, so the series should not be treated as perfectly comparable year to year. The 2025 survey involved 1,993 participants and was fielded June 25–July 29, 2025. McKinsey’s chart and survey details

The distinction corrects a figure sometimes repeated in the essay’s framing: McKinsey’s chart gives 72% for 2024, not 78%. Even accurately reported adoption figures say how many respondents reported organizational use under the survey’s definition; they do not show whether deployed models learn from that use.

How to assess claims about AGI progress

  • Ask what was measured. A task-duration benchmark measures success on tasks of a particular length and reliability, within its task suite.
  • Separate sustained action from learning. Look for evidence that experience changes capability, that the change persists, and that it transfers beyond the original context.
  • Check the domain and reliability limits. A result on software tasks does not automatically apply to other fields, and very long-duration estimates may be unreliable.
  • Ask who controls updates. Identify who supplies learning data, authorizes changes, evaluates outcomes, and can reverse an update.
  • Treat company proposals as proposals. For point-of-use learning, ask what is retained, how changes are tested, and how users can audit or undo them.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.