What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI progress should not be measured only by whether large language models (LLMs) become larger or better at generating text. Matt Asay’s 2024 InfoWorld analysis argues for a broader research portfolio—including reinforcement learning, recurrent neural networks, and diffusion models—while later work illustrates how LLMs can also be combined with tools and specialized components. His case for diversity is an opinion, not a settled scientific finding, but it raises a useful question: which approach fits the task, and what evidence shows it works?

What Asay means by looking beyond LLMs

In his 8 April 2024 InfoWorld analysis, Matt Asay challenges the idea that advances in AI are synonymous with advances in LLMs. He characterizes LLMs as systems that generate plausible text without grasping fundamental truth, and argues that their strengths are concentrated in statistical text tasks. He also contends that making models larger may bring only marginal gains on tasks outside text, and that LLMs should not be treated as a guaranteed route to artificial general intelligence (AGI). These are Asay’s interpretations, not conclusions established by a comparative scientific review. Read Asay’s analysis at InfoWorld.

The practical point is not that LLMs are useless or that another method has already displaced them. It is that AI encompasses different tasks, learning methods, and output types. A research agenda centered on one model family could miss capabilities that require other techniques—or combinations of techniques.

Approaches the essay says deserve attention

Reinforcement learning

Asay points to Diffblue’s Java unit-test generation as an example of AI work he describes as not using an LLM. He also makes a performance comparison for the system; that comparison is an assertion in the essay and is not independently verified here. The example is best read as evidence that Asay wants alternatives explored, not as proof that reinforcement learning is broadly superior for programming tasks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diffusion models

Asay cites Midjourney as an example of generative AI that does not depend on an LLM. The example highlights a basic distinction: generating images and generating text are different tasks, and an LLM is not the only possible foundation for generative systems.

Architectures can change what is possible

The essay invokes recurrent neural networks in the history of image recognition and transformers in text prediction as examples of architectural shifts that helped change capability. That is Asay’s framing of the history, not a comprehensive account of either field. Its broader implication is that researchers should remain open to architectures and methods that do not fit the current dominant pattern.

Why a broader portfolio is not an argument against LLMs

Thinking beyond LLMs does not require choosing between LLMs and everything else. A later example, the 2026 paper Accelerating scientific discovery with Co-Scientist, describes a Gemini-based multi-agent system for generating scientific hypotheses. It combines an LLM with specialized agents, web search and other tools, persistent context, iterative hypothesis review, and feedback from scientists. The system is still LLM-based, but its design distributes work across components rather than relying on a single model operating alone. Read the Co-Scientist paper.

This hybrid design illustrates one way AI research can move beyond a single architecture without abandoning LLMs: use a model where it is useful, then add tools, memory, specialist components, or human judgment where the task calls for them. One system, however, cannot settle which approach will produce the next major advance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to assess claims about AI progress

Comparisons are meaningful only when they specify what is being compared and what evidence supports the result. Useful questions include:

  • What is the task and output? Text generation, image generation, software testing, and scientific hypothesis generation are different problems.
  • How does the system learn or operate? Does it predict patterns, learn through interaction, or coordinate specialized components?
  • What supporting components does it use? Tools, persistent context, iterative review, and expert feedback can affect a system’s capabilities.
  • What kind of evaluation supports the claim? A benchmark score, expert assessment, and validation through real-world experiments answer different questions.
  • How broad is the evidence? Results from one task or system do not establish performance across AI as a whole.

The Co-Scientist paper reports automated evaluation across 203 research goals, including a subset of 15 expert-curated biomedical goals, and human expert evaluation across 11 goals. It also reports experimental validation in three biomedical application areas: drug repurposing, treatment-target discovery, and investigation of antimicrobial-resistance mechanisms. These counts describe the scope of that study, not general AI capability or proof that LLM-based systems outperform other approaches. The paper notes that some evaluations are small-scale and that expert ratings are subjective rather than objective ground truth.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the market-concentration argument does—and does not—show

Asay also warns that concentrated investment in LLMs could distort the AI market and crowd out other approaches. He attributes a related concern about market concentration to Tim O’Reilly. These are arguments about incentives and the research ecosystem, not quantified findings established by the sources discussed here. They support asking whether funding and attention are too concentrated; they do not establish how much concentration exists or what its effects have been.

What readers can take from the argument

  • LLMs are an important part of AI, but they are not the only model family or system design worth investigating.
  • Examples such as Diffblue and Midjourney show the range of approaches Asay wants readers to consider; they do not establish that those approaches will lead the next breakthrough.
  • Hybrid systems offer another path: LLMs can work alongside tools, specialist agents, iterative processes, and human feedback.
  • Claims about progress should be judged by task-specific evidence and the scope and limits of the evaluation—not by a model label alone.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.