Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Harness Score measures repository-level scaffolding around an AI coding agent: its instructions, tools, feedback mechanisms and guardrails. It reports an L0–L4 maturity level, a score out of 108 across six dimensions, and prioritized fixes. Use it to find structural gaps—not as proof that an agent writes correct or safe code.
What Harness Score measures
An AI coding harness is the context and control system around a model: repository guidance, available tools, feedback and safeguards. Harness Score scans selected repository artifacts to assess part of that setup. The project README describes 36 filesystem-based checks; the scanner does not use LLM judgments or network lookups. Its results include a maturity level, dimension scores and remediation guidance. Harness Score’s README documents its current model and usage.
The project’s README, accessed in 2026, assigns up to 108 points across six dimensions:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors| Dimension | Maximum points |
|---|---|
| Context & Guides | 20 |
| Skills & Commands | 17 |
| Hooks & Guardrails | 14 |
| Sensors & Feedback | 20 |
| CI Feedback | 14 |
| Hygiene & Safety | 23 |
| Total | 108 across 36 checks |
These are Harness Score’s product-specific figures, not an independently validated reliability metric. The project notes that checks, point totals and level thresholds can change in minor releases, so record the version or access date when reporting a result.
#1 Best Overall
What the L0–L4 levels mean
The project defines its own maturity ladder. It is not an industry-wide standard, and levels are not simply score ranges: reaching a higher level requires covering new kinds of repository support, not merely accumulating points.
L0 · Unharnessed
There is little structured repository guidance for an agent. The project recommends starting with an AGENTS.md file.
Rank #2
L1 · Documented
A substantive AGENTS.md orients an agent to the project, its build and test process, and relevant constraints.
L2 · Guided
Guidance becomes more specific through scoped rules and at least one skill or command, with basic hygiene. The guidance is versioned alongside the code.
L3 · Sensing
Tests, linting, type checking and CI provide repeatable feedback on pushes, giving the agent workflow and its maintainers ways to detect problems.
L4 · Self-correcting
Runtime gates and feedback hooks close more of the loop: they can block risky actions and apply checks such as linting or formatting inline.
Rank #4
How to use the score to improve a repository
- Run the scanner on the repository you want to assess. Follow the npm installation and CLI instructions in the project README; the project documents Markdown, machine-readable and badge output.
- Read the level and dimension breakdown together. The overall level describes the project’s maturity classification, while the six dimension scores show where repository scaffolding is present or missing.
- Inspect the specific checks blocking the next level. Choose fixes that address meaningful gaps in context, tools, feedback or guardrails rather than trying to optimize points in isolation.
- Make a change, then run the scan again. Comparing results helps confirm whether the repository artifacts changed as intended. Compare like versions, since model details may evolve.
- Automate a threshold if it fits your workflow. The project documents a GitHub Action and a minimum-level CI gate. A passing gate means the configured structural threshold was met; it does not certify that the repository or agent is safe.
What a high score cannot establish
A high score means relevant infrastructure exists, not that it works well. The project explicitly says the scanner does not assess test quality, whether rules are true or current, functional correctness, or team practices such as review culture and branch protection. Those matters require evidence beyond a filesystem scan.
Free tools Windows power users keep installed
One-click scans. No signup required.
For a fair comparison between repositories, hold the Harness Score version and repository scope constant. Compare the overall level, all six dimension scores, failed checks and their remediations, then assess agent behavior separately on a fixed set of tasks. A raw score alone can mislead when repositories have different purposes or constraints.
Best Value
Pair structural scanning with behavioral evaluation
A repository can contain tests and hooks without those controls catching the failures that matter. Behavioral checks address a different question: what does the agent actually do when it encounters a situation? Google’s engineering guidance recommends observable behavioral checks as an iteration aid and treats them as complementary to end-to-end benchmarks. For example, a team might check whether an agent asks for clarification when a request is ambiguous, runs a validator after changing a build file, or uses only an allowed tool. These are examples to adapt, not a universal required suite.
Google’s September 9, 2026 guidance says behavioral checks can observe intermediate actions, while end-to-end benchmarks measure final task outcomes but usually do not explain why a score changed. For noisy model behavior, it recommends evaluating batches and looking at aggregate trends rather than relying on one run. The authors write: “A robust harness evaluation framework separates behavioral assertions into fast, deterministic, unit-style checks that run locally.” Google Developers Blog, September 9, 2026.
Do not confuse Harness Score with other uses of “harness”
The word “harness” also refers more broadly to runtime scaffolding that drives model and tool calls, manages state and context, applies approvals, and supports multistep work. Microsoft Learn describes runtime components such as chat pipelines, context providers, middleware, observability and optional bounded loops. That general runtime concept is broader than Harness Score’s repository-artifact scanner. Microsoft Learn’s agent-harness overview explains the runtime usage.
Harness Protocol is a separate portability proposal for a vendor-neutral harness.yaml describing plugins, MCP servers, environment requirements, instructions and permissions. Its documentation describes schema v1 as current, with exchange and registry layers planned. It is not the Harness Score L0–L4 scale. Harness Protocol documentation.
A 2026 arXiv preprint proposes a different H0–H3 controlled-visibility ladder and trace-based evaluation approach. It is separate research, not the Harness Score maturity scale or an adopted standard. The preprint.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

