Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In one developer’s 2026 experiment, Strands, LangGraph, and CrewAI all passed a narrow four-key JSON task, but differed on approval handling and how much tool activity could be reconstructed from traces. The tests also exposed a sharp failure mode: Strands produced three empty final answers despite successful process exits. These are measurements from one setup—not a general framework ranking or evidence of production reliability.

What the 45-run experiment measured

Developer sunnydachs reported 45 runs across three task groups: 18 approval-gate runs, 36 audit-trail runs analyzed, and nine structured-output runs. These are the experiment’s reported counts; they do not mean 45 runs per framework or establish population-level failure rates.

The same model, tools, and recorder proxy were used, according to the author. The approval task created a news digest, sought approval, and then attempted a simulated publish action. The audit task scored whether seven audit-relevant facts could be recovered from traces, including rationale, tool order, tool arguments, and model identity. The structured-output task required exactly four JSON keys: summary, word_count, topics, and publish_ready.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The author did not name the model/provider or framework versions. The linked experiment repository is described as containing commands, but the results here are the author’s reported measurements, not independently reproduced findings. Read the author’s full account.

How approval behaved in the tested workflows

The approval implementations differed in where control was placed: in the model’s instructions, in graph control flow, or in a task configured to request human input. Results below describe only the author’s simulated workflow.

Framework Approval mechanism and reported result
Strands The prompt asked the model to seek reviewer approval. The author reports correct ordering in 3 of 3 runs and no publish after rejection, but one run called the simulated publish action twice.
LangGraph interrupt() paused the graph and Command(resume=...) continued it. The graph suspended in all 6 reported runs; on rejection, routing avoided the publish action.
CrewAI Task(human_input=True) requested console feedback after the task. Approval took one call and rejection took two in the described test. Repeating the same rejection led to a reported loop of 131 LLM calls.

For a workflow where approval must be an enforced gate, these observations make it important to test whether the workflow itself blocks the side effect or merely asks the model to wait. The duplicate publish call and repeated-rejection loop also show why approval tests should track side effects and call counts, not just whether the model acknowledged feedback.

What the traces did—and did not—show

In the author’s trace-only reconstruction scoring, rationale was recoverable in 100% of runs for each framework. Tool order and arguments were reported as recoverable in all Strands runs, none of the LangGraph runs, and half of CrewAI runs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Framework Rationale recoverable Tool order recoverable Tool arguments recoverable
Strands 100% 100% 100%
LangGraph 100% 0% 0%
CrewAI 100% 50% 50%

The author attributes LangGraph’s missing tool-order and argument evidence to tools being called in code rather than appearing as model tool calls on the wire in this setup. That is a trace-reconstruction result, not proof that LangGraph cannot provide audit evidence through other instrumentation. Nor do these percentages establish regulatory or overall audit compliance.

Strict JSON passed, but empty results matter

All three frameworks met the tested exact four-key JSON requirement in all nine reported structured-output runs. The reported word count matched the summary length each time. For Strands, the author says the validation loop averaged two calls, with four revisions in one run.

Across the broader 45-run set, Strands had three empty final outputs that nevertheless exited successfully. LangGraph and CrewAI had no such empty outputs in the reported tests. The author says the completed result could be recovered from a tool-call argument in the trace, but a downstream consumer could still receive an empty final answer. A successful process exit should therefore not be treated as a guarantee that a usable deliverable exists.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to use these findings in a framework evaluation

The benchmark is most useful as a source of test cases, not as a scorecard for choosing a universal winner. The author summarized the limits as: “One model, 3 runs per cell – directional, not a definitive ranking.” One model, a scripted human, three runs per cell, no real notification or user-interface flow, and a simulated destructive action constrain what the results can establish.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For your own evaluation, make the workflow and its failure conditions explicit. Include checks such as:

  • Approval enforcement: Can an unapproved or rejected request reach the side-effecting action? Does the gate live in control flow or depend on model compliance?
  • Duplicate effects: Can retries, resumed runs, or repeated feedback invoke the same action more than once?
  • Trace retrieval: Can an auditor reconstruct the decision rationale, tool sequence, arguments, and model identity from the records your deployment actually stores?
  • Output guarantees: Does a successful exit also produce a nonempty result that validates against the required schema?
  • Recovery: What happens after a process crash during approval, tool execution, or final-output generation?

The reported tests do not answer how these frameworks behave across other models, versions, interfaces, or production workloads. Measure those conditions directly before making a deployment decision.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.