Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An empty architecture baseline after a 121-file refactoring sounds like a clean result. According to the author’s account, it was not a merge-ready result. Alexander Kell reports that review still found two merge-blocking failures: a public Python API regression, and a shell gate that could stay green after a failing test command. The architecture check passed, and the code still was not ready to ship.

What the experiment reported

Kell describes running ArchKeel, an architecture-boundary tool, through a 121-file refactoring experiment on DATAMIMIC CE. The key claim is that the target architecture was written down before any coding agent touched the code. He reports that the target was not widened during the work to make the numbers look better.

The reported figures are all the author’s own account. None of them comes from an independent audit or measurement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Figure Reported value Who reported it and what that means
Files in the refactoring experiment 121 Reported by Alexander Kell. No independent auditor is named.
Declared architecture violations at the start 613 Reported by Alexander Kell as the starting count. No independent measurement is cited.
Steps to reach the final state 11 Reported by Alexander Kell. The steps themselves are not described in the summary available.
Elapsed time Roughly 6.5 hours Reported by Alexander Kell as approximate.
Final architecture baseline Empty Reported by Alexander Kell as the outcome of the architecture check.

The article is listed on DEV Community as a six-minute post by “Alex,” dated Sep 22. The listing extract does not show a year, and the full body was not accessible for verification, so the detailed procedure, repository state, software versions and agent configuration are not established here.

#1 Best Overall
Inspiration Software, Inc.
  • The premier tool to develop ideas and organize thinking...brainstorming, webbing, diagramming,
  • planning, critical thinking, concept mapping etc.

What an empty baseline does and does not show

An empty baseline means the declared rules were satisfied. It says nothing about whether the code as a whole behaves correctly. Kell’s own review shows the gap. The architecture check was a narrow instrument, and the failures he found sat outside what it measured.

Kell says the dependency contract “worked as specified.” His criticism is about coverage. The contract did not cover enough component APIs, the package layout, or internal complexity. A dependency rule can be fully satisfied while a public interface quietly changes shape, or while a module grows harder to read.

The two merge-blocking failures

Kell reports that review found the following, in his words: “Review still found two merge-blocking failures: A public Python API regression. And a shell gate that could stay green after a failing test command.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A public Python API regression

The first failure is a regression in a public Python API. A refactor can keep every declared dependency edge legal and still change what external callers see. An architecture baseline that checks which modules may import which others will not detect a changed function signature or a removed name in a public module. Catching that class of problem needs a check on the public surface itself, such as tests that exercise public imports and call signatures, or an API comparison against the previous release.

A shell gate that stayed green after a failing test

The second failure is more instructive for anyone who relies on pipeline status. A gate that reports success when the test command has failed gives false confidence, and it is worse than having no gate, because people stop looking. The summary Kell published does not explain the exact shell mechanism behind the bug. The usual suspects are a pipeline that discards the test command’s exit status, for example by piping its output through another command, or a wrapper script that exits with the status of its last line instead of the failing step. Those are general shell behaviours, not findings from the report, so treat them as hypotheses to check in your own scripts rather than as the cause in this experiment.

A practical test is to make the test command fail on purpose, such as with a deliberately failing assertion on a throwaway branch, and confirm that the gate turns red. A gate that has never been seen failing has not been shown to work.

Who owned the weak target

Kell also revises his own earlier criticism. He had blamed the coding agents for creating large re-export facades, which are modules that re-export names from other modules to present a stable public surface. He now says the implementation brief had explicitly asked for those facades, so the agents were following instructions. He writes: “The weak target was mine.” The lesson he draws is that the target architecture was the weak point, not the agents’ execution against it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Two separate questions to ask about any agent-assisted refactor

The experiment suggests keeping two questions apart. They are editorial axes drawn from the reported outcomes, not measured comparisons.

  • Is the declared architecture satisfied? ArchKeel’s baseline answers this. In Kell’s run it reached zero declared violations.
  • Do the merge gates actually run and fail when they should? This covers public API checks, test commands and any shell wrappers around them. In Kell’s run, two gates were not sufficient on their own.

An architecture check can pass on both axes and still leave the refactor unfinished. A green baseline is a statement about one set of rules, not about the release.

What is not established

  • The full article body was not available for verification, so the step-by-step procedure, the repository, the agent models and the test suite are not described here.
  • No independent study or third-party measurement of these results was found. The numbers are one author’s report of one experiment.
  • The summary does not say how the two failures were fixed, or whether the fixes were re-verified.
  • The LinkedIn post and its short link to the full article could not be checked directly, so the quotations above come from the published summary.

Read the experiment as a useful warning about where architecture checks stop, not as evidence about how well coding agents perform in general.

Quick Recap

Bestseller No. 1
Inspiration Software, Inc.
Inspiration Software, Inc.
planning, critical thinking, concept mapping etc.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.