Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A coding agent can leave every test passing and still make a codebase harder to maintain. It may put a helper in the wrong module, import a database client into a domain layer, or replace statically resolvable code with indirection the analyzer cannot follow. A test suite checks the behaviors someone thought to test. It does not check whether component boundaries still hold. The approach described by Archkeel, an open-source tool whose author documented it in a DEV Community post in 2026, adds a separate architecture gate that runs alongside tests and rejects changes that weaken either the declared architecture or the evidence used to check it.

Why passing tests say little about architecture

Tests answer one question: did the exercised behaviors produce the expected results? They are silent on questions such as whether an application service now reaches directly into a persistence adapter, or whether a utility function has drifted into a module that other layers depend on. The author of the Archkeel write-up describes exactly this pattern in agent-assisted work: utilities placed in unsuitable modules, public interfaces crossed, and clients imported into layers that were never meant to know about them.

The gap is structural. A coding agent optimizes for a change that compiles, passes the suite, and satisfies the prompt. Nothing in that loop forces the agent to ask whether the dependency it just added is one the architecture allows. Architecture tests written with tools such as ArchUnit, import-linter, or dependency-cruiser close part of that gap, but the author’s point is that a check which only inspects the final snapshot misses a second problem: the analyzer’s view of the code can get weaker while the snapshot still looks clean.

What the gate checks

The gate models a target architecture as a contract. The contract lists components, the packages each component owns, the public names each component exposes, and explicit dependency rules between components. Every ordered pair of components must receive a decision, either allowed or forbidden, with a written reason. A pair with no decision stays open, and validation stays red until someone resolves it. The contract is therefore a complete map of the relationships the architect has thought about, not a list of prohibitions that happens to be incomplete.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The gate keeps three verdicts separate rather than folding them into one score:

Verdict Question it answers What a failure means
observation_complete Did the scan see everything it claims to see? The analyzer lost visibility, so other verdicts cannot be fully trusted.
declared_rules Does the code obey the contract? A declared dependency rule is broken, such as a forbidden import or cycle.
expectation_fulfilled Did the change match what was declared beforehand, without regressions? The change did something its published expectation did not describe, or made the evidence weaker.

Keeping these apart matters. A change can pass declared_rules cleanly while observation_complete has fallen, because the analyzer now sees less of the program. Without a separate verdict for visibility, that clean rule check would look like a clean result.

Who decides the rules

The architect remains responsible for the intended target architecture. The described tooling supports two modes. In interview mode, the packaged skill prepares recommendations from existing architecture documents and asks about conflicts and gaps. In auto mode, the skill makes the allow and forbid decisions itself and labels which party decided each rule, so a reviewer can see which decisions were machine-proposed.

Visibility loss as a regression

The most instructive part of the approach is the treatment of weaker evidence. In the author’s fixture, a change replaces two statically resolved calls with a dictionary lookup. The tests still pass. No forbidden import appears and no cycle is introduced. But the analyzer now reports one unresolved call where it previously reported none. A gate that only checks rules would accept this. The described gate treats the loss of static resolution as a regression and rejects a change that did not declare it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unresolved calls are compared by integer cross-multiplication rather than rounded percentages, so a small change in the ratio cannot disappear into rounding. The author states that unresolved calls are counted and reported, not estimated or guessed.

Process evidence: expectation before implementation

Rules describe the code as it is. The expectation mechanism describes the change the agent intended to make, and it has to exist before the implementation does. The sequence the author describes works like this:

  1. The agent commits an expectation file that states the intended architecture change: which components are affected, what relationships should appear or disappear, and what should not change.
  2. The agent submits the implementation in a separate, later commit or merge request.
  3. The gate checks Git ancestry and, where available, host merge request history to confirm that the expectation was published before the implementation.
  4. The gate compares the actual change with the declared expectation and reports any difference as a finding.

An expectation written after the fact cannot satisfy this check, which is the point. An agent that describes its change after seeing the diff can always make the description match the diff.

Exit codes and unknown evidence

The author’s stated exit codes are 0 for a pass, 1 for a rejection, and 2 when the input cannot be verified. The design rule is that unknown evidence never becomes a pass. The author summarizes it as “Unknown never becomes green.” A missing expectation, an unreadable history, or an analyzer that cannot run returns 2, not a silent success.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reported results and how far they reach

The figures below come from the author’s own account of applying the tool to a field-service application and to Archkeel itself. They are author-reported measurements from one project, not independent benchmarks, and none of them should be read as general accuracy for the tool.

Figure What it measures Scope stated by the author
140 of 156 component-pair decisions matched (89.7%) Agreement between the tool’s auto-mode decisions and the expected decisions One service, measured once. The author says it is not a general accuracy estimate for auto mode.
13 components Size of the field-service application’s target architecture The single application used in the account
162 violations in the first report against the final target Rule violations found on the first run, including 148 on the use-case-to-persistence-adapter dependency The field-service example only
630 unresolved calls out of 3,303 (Archkeel itself) and 998 out of 4,318 (the service) Calls the analyzer could not resolve statically Counted and reported, not estimated, for these two codebases
6 components, 30 component pairs, 46 rules in Archkeel’s self-check contract Size of the tool’s own architecture contract The author reports planting a violation to confirm that each enforcing rule fires

The field-service example reportedly used Python 3.12, FastAPI, async SQLAlchemy, PostgreSQL with PostGIS, Redis, Taskiq, and OR-Tools. These describe the environment in that account. They are not requirements for using the approach.

What the gate does not cover

The limits are as important as the design. The author states the following:

  • Runtime behavior, data flow, and performance are not observed. The gate sees static structure, not what happens when the code runs.
  • Two competing implementations of the same idea are not detected unless a rule or regression exposes them.
  • Private access through a package import, such as import pkg; pkg._member, can slip through.
  • Publication-order evidence does not prove that no one edited privately before publishing. It proves the order of what was published, not the history of what happened on a laptop.
  • A decision reason is checked for existence, not for truth. The tool confirms that a reason is written; it cannot confirm that the reason is correct.
  • Host evidence is GitLab-only at publication. The author describes GitLab merge requests as the source of host history, with no GitHub adapter at that time.

Determinism testing

The author tested determinism by running reports repeatedly across two clones, varying paths, hash seeds, working directories, time zones, and locales. Output was reported as byte-identical on one machine and one Python build. Cross-platform and cross-version determinism has not been established by that account, so teams running on other operating systems or Python versions should verify output themselves before relying on exact-match comparisons.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How this differs from snapshot rule checks

Dimension Snapshot rule checks Gate described by the author
Comparison basis Current code against declared rules Baseline against candidate, so weakened evidence is visible
Analyzer visibility Usually assumed complete Reported as its own verdict, observation_complete
Intent checking Not part of the check Change must match a published expectation committed earlier
Output Typically pass or fail per rule Three separate verdicts plus diagnostics, no aggregate score
Coverage Static dependencies and rules Static structure only; runtime, data flow, and performance not observed
Host dependence Generally none beyond the code GitLab merge request history used for publication order

The gate supplements tests and review. It does not replace them, and it does not replace human ownership of the architecture. The author presents it as a guardrail with explicit boundaries, not a verifier of design quality.

Getting started

The author describes Archkeel as MIT-licensed and distributed through GitHub and PyPI. A minimal first step is to inspect the command-line interface before adopting any contract:

  1. Run uvx archkeel --help to confirm the tool installs and to list its commands.
  2. Write or generate a contract that names your components, their owned packages, public names, and an allow or forbid decision with a reason for every ordered pair.
  3. Resolve every open pair before treating validation as meaningful, since open pairs keep the result red.
  4. Commit an expectation before the next agent-written change, and let the gate compare the change against it.

Command behavior and packaging can change between releases, so confirm the current options against the project’s own documentation before wiring the gate into a pipeline.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where the approach fits

The approach is most useful where agents make many small structural decisions across a codebase that already has a stated architecture: layered services, adapters around persistence and external clients, or domain packages that must not depend on infrastructure. It is least useful where no target architecture has been written down, because the contract then has nothing to enforce. In that case the first work is deciding the architecture, which no tool does for you.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Teams on GitLab can test the publication-order check directly. Teams on GitHub should expect to supply their own equivalent of the merge request evidence until an adapter exists.

The central idea is not specific to one tool. A green test suite shows that tested behavior survived a change. A gate that also checks what the change was supposed to do, and whether the analyzer could still see it, answers a different and harder question. In the author’s account, that second question is where the architecture quietly gets worse.

Alex, the article’s author, puts the principle plainly: “A gate that an agent can talk its way around isn’t a gate.”

The Bottom Line

Passing tests confirm the behaviors you tested, not the architecture you intended. Pair them with a contract that forbids specific dependencies, a check that the analyzer still sees the code, and an expectation committed before each change. Treat any unverifiable input as a failure, and keep the reported results in mind as one team’s account rather than proof of general accuracy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.