Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI agents can take on product and software work without leaving people out of the decisions. In Shinsuke Kagawa’s account of building an agent team, scripts handled coordination, agents debated product choices and reviewed implementation, and a human still set goals, accepted phases, and changed the team’s rules. The key lesson is not that this setup is proven to work everywhere; it is that delegation depends on deliberate process and clear human approval points.

What Kagawa’s team did

Kagawa describes a six-member team: an orchestrator, two directors using different model families, and three executors responsible for implementation, overflow work, and images. Agents did not message each other directly. Instead, a board of Markdown files carried requests and status.

A Bash script called studio polled the board every 30 seconds, launched agents for assigned work, routed tasks through the orchestrator, and carried messages between the board and Slack. While work was active, it ran a review patrol every 30 minutes. Scripts also managed waiting, handoffs, and crashed runs, reducing reliance on an agent remembering operational details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How product decisions were made

The decision process was designed to keep initial ideas independent and make conclusions checkable. Kagawa’s four stages were:

  1. Agree on the user goal independently. Directors first identified the goal without seeing one another’s framing.
  2. Generate proposals in fresh sessions. They wrote candidate ideas before examining the repository.
  3. Discuss unresolved questions. The directors continued until the open questions had answers.
  4. Check the conclusion. A director who had not authored the conclusion checked whether it met the goal, answered questions, supported changed views with evidence, compared alternatives, and tested inherited product choices against other options.

This ordering aimed to separate the goal and candidate ideas from the constraints and assumptions embedded in an existing codebase.

What went wrong in the first debate process

Simultaneous revisions produced unproductive switching

In the first cross-review design, directors revised their votes at the same time without seeing the other’s revised choice. A “yield on taste” rule led both to yield and repeatedly switch sides. Kagawa replaced that procedure with direct, turn-taking discussion.

In six discussions, the director who spoke first wrote the conclusion in four; three discussions ended on the first turn with simple agreement. These small, author-reported counts are observations about this workflow, not general success rates. Kagawa later randomized the first turn using a hash of the item name and required both directors to take at least one turn before a conclusion could be reached.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The orchestrator’s framing influenced decisions

Kagawa found that the orchestrator’s paraphrases and suggested options could frame what directors chose. The revised brief used the original request and Kagawa’s full words as quotations rather than relying on the orchestrator’s summaries.

Repository context could narrow the ideas too soon

In one example, directors suggested a narrow product adjustment even though feedback said the product lacked distinctiveness. After repeated attempts, they reused an example Kagawa had offered only to illustrate abstraction. The process changed so directors extracted the goal first, created candidates in fresh sessions before reading the repository, and only then considered feasibility and cost.

How the team addressed leakage and checked decisions

An initial rerun surfaced an earlier conclusion because an agent found it on the shared board. For the next attempt, Kagawa limited the brief to one item and withheld other files until a proposal had been written.

In his retest, the concept changed its product type, subject, and main interaction. Discussion took six turns instead of two; one director changed position and named the example that changed it, and the check passed. This is an anecdotal before-and-after account, not an independent evaluation of the process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How implementation work was reviewed

During active building, a reviewer patrol inspected changes every 30 minutes for unnecessary additions or removals and work aimed at cases that would not occur. Kagawa reports 90 patrols, half of which found nothing. The others identified issues ranging from bugs to inconsistencies between decision records and code. These are his counts, not independently audited metrics.

When implementation fell short, a task could be escalated from the regular executor to a stronger executor. Kagawa reports seven escalations. In two cited cases, a reviewer noticed that parts of earlier agreements had been dropped; the escalated executor restored them.

Where the human remained in control

This was delegation, not human withdrawal. Kagawa retained responsibility for setting phase goals and stopping points, providing views, answering questions that required human input, accepting completed phases, and changing the team’s rules. As he described his experiment: “I design how the team works in some detail, but I stopped telling the directors what to decide about the product.”

At the time of writing, Kagawa says 11 agreements had passed through the check; three were returned for a specific correction before passing. One check was repeated with director names anonymized and returned the same result. Those figures describe this system and its limited sample; they do not establish how a similar process would perform elsewhere.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the broader research can—and cannot—say

Related studies offer context for why process choices may matter, but they do not test Kagawa’s workflow. Lou and Sun’s 2024 paper, revised in December 2024, reports that large language models were sensitive to biased hints in its experiments; chain-of-thought, reflection, and explicit instructions to ignore anchor hints were not sufficient mitigations in those experiments. Read the paper.

Choi, Zhu, and Li’s paper, submitted in October 2025 and revised in April 2026, describes identity-driven self-bias and peer deference in multi-agent debate, reporting peer sycophancy more commonly than self-bias in its experiments. Read the paper.

Sharma and colleagues’ paper, submitted in 2023 and revised in May 2025, reports sycophancy across five assistants and four free-form generation tasks. It also reports that responses matching a user’s views were more likely to be preferred in the analyzed preference data. Read the paper.

These findings help explain risks such as anchoring and agreement-seeking; they do not show that Kagawa’s ordering, turn-taking, or review checks reliably prevent them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Design choices to consider before delegating work

Kagawa’s case suggests questions to resolve when designing an agent workflow. His account does not establish that one answer is universally best.

  • Coordination: use a centralized board or allow direct agent-to-agent communication?
  • Idea generation: ask for independent proposals in parallel or move into sequential debate?
  • Context: expose the repository early or delay it until candidates are written?
  • Review: rely on the agents doing the work or assign a separate checker?
  • Authority: require human approval at phase gates or permit broader autonomy?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.