Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Sentinel-IR turns selected JavaScript code structures into machine-readable facts so an AI agent can investigate security questions without repeatedly receiving entire source files. In one benchmark reported by its author, Sentinel-IR paired with raw-source fallback answered all 87 test questions correctly using 71.3% fewer input tokens than raw source alone. That is a promising result from a limited, unreplicated test—not proof that the format is universally more accurate or cheaper.

What Sentinel-IR is—and what it is not

Sentinel-IR is a compact representation of security-relevant facts extracted from JavaScript syntax. Its author, jackymenCZ, describes it as a deterministic “fact layer,” not a programming language developers write. Instead of asking an agent to inspect all the code for every question, the approach gives it structured information about selected code behavior, such as routes, imports, exports, environment-variable reads, calls, and risk signals.

That can help an agent answer questions such as “does this merge request touch the network?” or whether a change adds a POST route that reads an environment secret. The representation is intended to make those facts easier to provide and inspect; it does not replace the source code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the fact layer is produced and used

The implementation described in the author’s article follows this path:

  1. JavaScript source is parsed into a tree-sitter abstract syntax tree.
  2. An AstFacts stage extracts selected details, including routes, exports, imports, environment variables, calls, and risk signals.
  3. Those facts are projected into Sentinel-IR, a sparse, flat representation that keeps non-empty arrays and enabled operations.
  4. An AI agent uses the representation to answer questions. If it cannot resolve a question from the facts, it falls back to raw source.
  5. The workflow proceeds to validation, simulation, and commit.

Risk signals retain evidence and line references, which can make a finding easier to trace back to code. The article describes extraction after parsing as local and deterministic, with no network access, model call, or I/O. Those are descriptions of the author’s implementation, not independently audited properties. See the original Sentinel-IR article for the implementation details.

Why raw-source fallback matters

Sentinel-IR’s current format omits empty categories. If there are no environment-variable reads or disk writes, for example, an omitted category does not reliably tell an agent that the answer is “none.” It may mean the fact is absent, or that the representation does not include an explicit empty set. In the author’s benchmark, five questions could not be resolved from IR alone because they asked about empty sets.

The hybrid workflow handles that uncertainty by consulting raw source when the structured facts are insufficient. This is a meaningful design choice: an agent should not treat a missing key as proof that a risky behavior is absent unless the format explicitly guarantees that interpretation. The benchmark’s fallback resolved those five questions, but the test does not establish how often fallback will be needed in other codebases or workflows.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the reported benchmark found

jackymenCZ reports testing 12 files with 87 questions and 267 actual LLM calls against gpt-6-astra. The figures below are the author’s results from that test, not independently reproduced measurements.

Input provided to the agent Input tokens Correct answers What the result means
Raw source 279,476 84 of 87 (96.6%) Baseline in the author’s test
Sentinel-IR only 58,549 82 of 87 (94.3%) Lower token use, but five questions were unresolved
Sentinel-IR with raw-source fallback 80,340 87 of 87 (100%) All test questions answered, using 71.3% fewer input tokens than raw source

The strongest comparison is the hybrid approach against raw source: in this one test, it answered all questions correctly while using 71.3% fewer input tokens. IR alone also used fewer tokens, but it did not match raw-source accuracy. The results support examining the combination of compact facts and fallback—not claiming that the IR by itself is more accurate.

File size changes the token trade-off

The author estimates a break-even point near 303 source tokens, or about 34 lines: below that rough size, encoding facts as IR can use more tokens than sending the source directly. The article’s examples include multiple small files with negative savings. Larger files more often showed substantial reductions, but the estimate is fitted to the author’s data and should not be treated as a universal cutoff.

For a real workflow, compare token use across the actual file mix, including the cost of raw-source fallback. A design that saves tokens on large files can lose some or all of that advantage if many files are short or unresolved questions trigger frequent source reads.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How much confidence to put in the results

The benchmark has important limits: it used one model, one run, and a small corpus owned by the author, with no variance analysis. Token counts for variants were estimated using characters divided by four; the author says the estimate came within 5% of provider billing for that run. This does not establish performance across other models, repositories, question sets, or repeated trials.

The author also reports validation across 16 external repositories and 140 merged pull requests. In that account, a critical gate blocked three pull requests involving external command execution; hand-verified findings had reported precision of 5/5 and recall of 85/85. These figures are author-reported validation, not an independent security evaluation, and they do not by themselves establish general detection rates.

The article reports a live-run cost of $4.93 on the organization account and says cache writes accounted for roughly 70% of cost in the benchmark setup. These are historical, setup-specific figures—not current pricing or a forecast for another user’s workload.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the Orbit Local comparison does—and does not—show

In a limited comparison reported by the author, GitLab Orbit Local answered 29 of 87 questions correctly (33.3%), had 41.4% context completeness, and gave seven confidently wrong answers. Sentinel-IR’s hybrid setup scored 87 of 87 (100%), with 100% context completeness and no confidently wrong answers in that comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is not an overall product ranking: the comparison is limited to Orbit Local and the author’s test setup. Orbit Remote was not measured because, according to the article, it required a Premium group and a Knowledge Graph: Read token. The results therefore say nothing about Orbit Remote’s performance.

How to assess Sentinel-IR for a code workflow

When evaluating a fact-layer approach, look beyond its smallest token count. Check whether it exposes the facts relevant to your security questions, preserves evidence and line references, distinguishes an empty result from an omitted category, and provides a dependable way to consult the source when facts are incomplete.

  • Answer quality: Test representative questions about routes, network calls, environment reads, writes, and spawned processes.
  • Unresolved cases: Record which questions require raw-source fallback, especially questions asking whether a behavior is absent.
  • Token use by file size: Include small files as well as large ones; the fact representation can be larger than the source for short files.
  • Traceability: Confirm that risk findings include useful evidence and line references.
  • Repeatability: Test more than one run and, where relevant, more than one model and repository before relying on a measured accuracy or savings figure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.