Yes, in a practical, output-focused sense: LLM agents can generate ideas or solutions that are novel and useful under particular tests. Whether that makes them “truly” creative in a human-like sense is a different, unsettled question. Studies disagree because they measure different tasks, kinds of novelty, and comparisons—not because one has established a universal verdict on creativity.
What does “truly creative” mean?
The answer depends on which part of creativity you mean. One definition focuses on the work itself: an idea or artifact is creative if it is sufficiently novel and useful, effective, or original according to stated criteria. Another asks how the work came about, including whether the process involves intention, personal experience, or social understanding. Success under the first definition does not settle the second.
Creativity judged by the result
Under an output-focused definition, an LLM agent can count as creative when its output clears a task’s criteria for novelty and usefulness. Those criteria need to be stated: an unusual answer is not necessarily a good one, and a useful answer may be conventional. In research on machine-learning engineering agents, for example, novelty is evaluated both against an agent’s own previous solutions and against human solutions, while usefulness is assessed through task performance.
Creativity judged by the process
A process-focused definition asks whether an agent has the kinds of intentions, experiences, or social grounding people associate with human creativity. A 2026 arXiv preprint, “On the Creativity of AI Agents,” argues that current agents show functional creativity while lacking key aspects of ontological creativity. That is the authors’ conceptual position, not an experimental result that settles whether an AI has inner experience.
Recommended Free Tools
#1 Best Overall
There is no universally accepted metric or test that resolves both definitions at once. A benchmark can show that an agent performed creatively on a particular task; it cannot, by itself, establish what the agent experiences or whether its agency is human-like.
Why do studies disagree about AI and human creativity?
“Creativity” covers tasks with different demands. Generating many possible uses for an object is not the same as writing a compelling story, devising an engineering solution, or working collaboratively. Results also depend on which model is tested, how prompts are written, how many responses are sampled, and whether evaluators prioritize originality, usefulness, or both.
Rank #2
| Study and setup | What it reported | How to interpret it |
|---|---|---|
| Wang et al., Nature Human Behaviour, published 23 December 2025; 9,198 human participants and 215,542 LLM observations on a divergent-creativity task | Average human creativity was slightly higher; humans showed greater variation and an advantage among the highest performers. Persona prompts helped only up to a threshold, while strategic prompt-engineering results were mixed to negative. | This is a large comparison on divergent idea generation, not a ranking of people and models across every creative field. |
| 2024 Scientific Reports study; GPT-4 compared with 151 people on the Alternative Uses Task, Consequences Task, and Divergent Associations Task | The authors reported higher GPT-4 scores on all three measures and greater originality and elaboration after controlling for fluency. | The result applies to GPT-4 on those tests and that sample; it does not establish that models generally outperform people at creativity. |
| 2025 Thinking Skills and Creativity paper; 13 creative tasks | The abstract reports LLMs averaging the 46th percentile against humans, with stronger performance in divergent thinking and problem solving than in creative writing. In the tested setup, ten repeated responses produced a collective result comparable to 8–10 people. | The result is specific to the tasks, models, and repeated-response setup described by the study; it is not a general equivalence between an LLM and a human group. |
| Microsoft Research report; 4,541 multi-agent LLM ideas and 341 human-team ideas across six problem-solving tasks | The report gives an effect size of Cohen’s d=1.50 and says the agents’ advantage was driven by novelty while usefulness remained comparable. It also reports that broader-ranging conversations were associated with more creative ideas. | This measures teams on six particular problem-solving tasks. It does not show that multi-agent systems are generally better creative collaborators. |
| Bhushan, Zhang, and Wang, arXiv preprint posted 30 August 2026; AIDE and AIRA-Dojo on ten Kaggle-style ML engineering tasks | The authors found that agents could explore novel solutions without converting that novelty into improved task performance; novelty declined as agents moved from exploration to exploitation. | This is evidence that novelty and success can diverge in a specific engineering setting, not a judgment about every kind of creative work. |
These findings are not interchangeable. A divergent-thinking score measures something different from creative-writing quality; an agent team is not the same unit of comparison as a single model response; and average performance can conceal a difference among top performers. The task and scoring method matter as much as the headline comparison.
Does producing something novel mean an agent is creative?
Not necessarily. Novelty means a result is new or uncommon relative to a reference point. Usefulness asks whether it works or serves the task. A system can produce unusual ideas that do not improve its results.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The ML engineering evaluation by Bhushan, Zhang, and Wang makes this distinction explicit. It considers psychological novelty within an agent’s own run, historical novelty relative to human solutions, and usefulness in completing the task. The authors report that agents sometimes reached solutions more novel than medal-winning human solutions while still performing worse on the task. In this setting, novelty was not a substitute for effectiveness.
That distinction is important outside engineering, too: originality is one ingredient in many accounts of creativity, but the evidence here does not support treating originality alone as proof of successful or meaningful creative work.
Can multiple agents—or repeated prompts—be more creative than one response?
They can change the output being evaluated, but the comparison depends on how those outputs are generated and scored. Repeated sampling may produce a stronger selection of ideas than a single answer, while multiple agents may contribute variety through interaction. Neither setup automatically demonstrates human-like collaboration or general superiority.
The 2025 13-task study’s reported comparison between ten repeated responses and 8–10 people is a result from its tested setup. Separately, Microsoft Research’s six-task study found multi-agent teams ahead of the human teams and single agents it evaluated, with novelty driving the advantage and usefulness remaining comparable. These are bounded findings, not evidence that a particular number of prompts or agents will reliably match a human group on any open-ended task.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Best Value
What can LLM agents usefully do in creative work?
The studies support treating agents as capable creative assistants for bounded tasks—not as automatic authorities on what is original, appropriate, or true. A person can set the goal, ask for alternatives, judge whether an idea fits its audience and constraints, verify factual claims, and revise the result. That division of work uses an agent’s capacity to generate possibilities without assuming that a plausible-sounding output is necessarily valuable.
- Use them to expand options: request several distinct approaches, then compare them against the brief rather than accepting the first fluent answer.
- Separate originality from quality: assess whether an idea is genuinely different and whether it solves the actual problem.
- Check claims and constraints: verify factual details and test whether a proposed solution works in the real context.
- Keep human judgment in the loop: choose, edit, or reject outputs based on goals, audience, and consequences that the task’s scoring criteria may not capture.
What is the most accurate conclusion?
LLM agents can meet task-based, output-focused tests of creativity, and some studies find strong performance on specific measures or team tasks. Other evidence finds slightly higher average human performance, a stronger human advantage among top performers, or novel agent solutions that do not translate into better task results. These conclusions can coexist because the studies test different things.
So the evidence supports a qualified answer: agents can produce novel and sometimes useful work, but no result described here establishes that they possess human-like creative agency or experience. “Creative on this task” is a defensible claim when the task and criteria are clear; “truly creative” requires deciding whether the output is enough or whether the process and its relation to human experience matter too.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

