iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A decision layer is the set of checks, routing choices, and judgments around an AI agent that decide what context the agent receives, which mechanism handles each choice, and whether a proposed action may run. Its practical test, as Sunil Ramlochan frames it in his September 29, 2026 article “The Decision Layer – A Practical Architecture for Building Cheaper, Safer AI Agents,” is to ask, for each decision, what the least expensive mechanism is that can make it reliably, given what happens if it is wrong. The answer is often ordinary code, sometimes a narrow model, and occasionally a person. Only rarely does it justify sending everything to a larger model.
What counts as part of the decision layer
The decision layer is not one model or one box in an architecture diagram. It is a collection of mechanisms that do four jobs: they shape incoming work and the context an agent sees, they choose among models or tools, they govern which actions may proceed, and they decide whether the evidence is sufficient to stop. Any of these can be implemented in code, in a model, or with a human in the loop. The useful question is not “where is the AI?” but “which mechanism is making this call, and what does it cost to get it wrong?”
Choosing a mechanism for each decision
The article’s central move is to stop treating an agent as a single decision-maker and to pick a mechanism per decision. Explicit, authoritative rules belong in ordinary code. Interpretation of messy text may suit a focused model. Difficult investigation may justify a stronger reasoning agent. Some decisions, especially those involving authority or context that software cannot see, belong to a person.
| Mechanism | Best suited to | Typical weakness | Example from the decision-layer framework |
|---|---|---|---|
| Ordinary code and rules | Decisions with explicit, verifiable criteria, such as arithmetic or scope checks | Brittle when the criteria are fuzzy or the inputs are unstructured | Keeping arithmetic such as totals in code rather than in a model |
| Focused model | Interpreting ambiguous text against a defined answer set | Can be confidently wrong; needs an evaluation set and an “uncertain” path | Ranking which candidate file to read early, with a narrow output set |
| Stronger reasoning agent | Open-ended investigation where the path is not known in advance | Highest cost and latency; the most room for unbounded actions | The coding agent that investigates a bug across the codebase |
| Person with authority and context | Decisions that depend on policy, accountability, or information outside the system | Slow, and a bottleneck if used for routine calls | Approving an action that changes data or money |
The cheapest mechanism that is reliable enough is the target. A stronger model used where a rule would do adds cost and new failure modes without adding control.
#1 Best Overall
The four places to inspect
The article identifies four decision points: before assembling context, before choosing among models, before an action executes, and after a result arrives. These are places to consider, not mandatory extra model calls. For some of them, ordinary search or a deterministic check will be the better tool. Treat the list as a map of where errors can enter, not as a prescription to add a decision step at each one.
Writing a contract before extracting a decision
Before any decision is moved out of a general prompt and into its own component, the article recommends writing a contract. Each contract should specify:
- The question, stated precisely enough that two reviewers would classify the same case the same way.
- The available evidence, meaning exactly the inputs the component will see.
- The answer set, including an “uncertain” or “insufficient evidence” option where appropriate.
- The acceptance rule, which defines how an answer will be judged correct.
- What the answer may influence, such as the order of reading or whether an action is blocked.
- The fallback or escalation when the answer is uncertain or the component fails.
- The consequence of error, in both directions.
Consider file selection for a coding agent. A workable answer set is “read early,” “read later,” or “insufficient evidence.” The contract must also say that “read later” does not mean the file is dropped from the investigation. If a classifier’s “read later” quietly removes a file, the system has made an exclusion decision that nobody designed or measured.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Asymmetric mistakes
The framework insists that errors are not equal, so a single accuracy figure hides the decision that matters. Two asymmetries recur in the article’s examples:
Rank #2
- Including an irrelevant file wastes some reading. Excluding the file that contains the cause can derail the whole investigation.
- A false block delays work. A false permission can damage data.
Measure each error type separately and attach its real cost. A component can look accurate overall and still be unacceptable if its rare errors are the expensive ones.
Keeping interpretation separate from authorization
The clearest rule in the article is that interpretation and authorization are different jobs. A model can help assess whether an action looks risky. Authority to act must come from independent software that enforces scope, access rights, and required approvals. Ramlochan puts the principle this way: “Models can interpret policy. Software should enforce policy.”
For agent permissions, the article gives an illustrative sequence:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors- The agent proposes an action.
- Scope and permission checks run in software.
- An additional risk assessment runs where it adds value.
- Required approval is obtained.
- The action executes.
Two failure modes follow from this. The approval requirement must not depend on the agent remembering to ask, and a prompt warning is not a permission system. A second model’s approval also does not independently establish authority. Each of those is a model judgment standing in for an enforced control.
Rank #3
Checking results against the requirement
After a result arrives, the article separates what direct evidence proves from what it does not. A successful edit proves that an edit occurred. Passing tests prove that those tests passed under their run conditions. Neither alone shows that the customer’s reported problem is fixed. The practical rule is to use direct system evidence where it exists, and then to interpret whether that evidence addresses the requirement. A reader-facing version of the same check is: “Does the test cover the reported behavior?”
A rollout order that earns added authority
The article’s recommended rollout is Contract, Shadow, Measure, Gate, Learn. In practice it runs as follows:
- Contract. Write the decision contract and record a baseline for the current workflow.
- Shadow. Let the new component predict without controlling the workflow.
- Measure. Review disagreements and track missed evidence, unnecessary inclusions, delay, total cost, rework, and task quality.
- Gate. Grant limited authority only when the evaluation supports it, and define timeout, invalid-output, uncertainty, and rollback behavior before that happens.
- Learn. Keep test examples separate from the examples used to tune the component, so that the measurements remain honest.
The author’s phrase for the discipline is “A good architecture earns its complexity.” Each added layer should be justified by a measured result, not by how sophisticated it looks.
Free tools Windows power users keep installed
One-click scans. No signup required.
Worked example: an incorrect checkout total
The article’s running example is a coding agent investigating a checkout total that differs from the expected amount. The proposed file-priority helper receives the bug report, the test output, a candidate file path, and bounded excerpts, then recommends which file to inspect early. The article uses a candidate question such as “Could this file help explain why the checkout total differs from the expected amount?” and shows Jev’s Choice interface and TypeSafe’s Playground as one way to prototype the bounded choice.
Rank #4
Three limits matter here. The checkout case, the training examples, and the operating policy are presented as a design, not as an experiment. The article states that the specialist was not trained or measured, so its sample classification of a conversion file is an example answer, not an observed model response. The article also reports that the referenced Jev 1.13 documentation lists risks involving numerical precision, indirect reasoning, and adversarial content. Those vendor materials were not independently checked for this piece, so confirm current behavior and documentation before relying on any version-specific claim. The article’s practical boundary is to keep arithmetic in code and let the coding agent do the investigation.
Cost is a full-task question
The framework rejects optimizing for the price of a single call. Compute, latency, rework, missed evidence, review effort, and error consequences all shape the real cost, and the test is whether quality and completion remain acceptable as the total changes. The article illustrates this with hypothetical figures that are not measurements from any organization or year:
| Line item (hypothetical, from the author’s example) | Amount |
|---|---|
| Original workflow | $1.00 |
| New decision layer | $0.08 |
| Remaining investigation | $0.65 |
| Average additional rework | $0.30 |
| New total, including rework | $1.03 |
The example’s own total sits slightly above the original, even though the new component itself is cheap. That is the point of measuring the full task: a cheaper step does not guarantee a cheaper outcome. The article does not report any measured saving, and none should be read into these figures.
Comparing candidate mechanisms
When two or more mechanisms are candidates for the same decision, compare them on the same axes:
Best Value
- The evidence available to each component.
- Decision quality, and the kinds of errors each one makes.
- The consequence and reversibility of those errors.
- Latency and compute.
- Full task cost, including retries, human review, and rework.
- The permission and authority boundary.
- Fallback behavior when the component is uncertain or unavailable.
- How easily the component can be inspected and replaced.
These are evaluation axes drawn from the framework. The article does not present a measured head-to-head comparison of mechanisms, so no ranking should be inferred from the list.
Where the framework stops
A decision layer is one part of a larger system. The evidence the component sees, the deterministic checks around it, the permissions that bound its effects, and the fallback when it fails together determine whether any decision can safely influence execution. The framework describes how to divide the work. It does not show that doing so makes agents cheaper or safer in general. Those claims would need the shadow-and-measure process above, run on the system in question.
The article also cites external guidance from Anthropic, OWASP, AWS, and others. Those sources were not independently verified for this piece, so treat the framework’s references to them as pointers to check rather than as settled facts.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

