Recommended Free Tools
An SRE agent that gets every test incident right may only be recalling incidents it has already seen. I learned that the hard way: my first 10/10 evaluation used cases already stored in the agent’s memory. It showed that the system could find relevant stored material—not that it could diagnose unseen incidents.
For a more meaningful test, I kept 104 of 114 OpenSRE incidents in a fresh Hindsight memory bank and held 10 out. I asked symptom-only questions about those held-out cases, then compared answers with memory against a baseline using the same model and prompt but no memory block. In that small evaluation, the memory-backed agent classified 9 of 10 correctly; the baseline had no fully correct answers. The result is promising, but it is not proof of general reliability.
Why the first 10/10 result was misleading
My initial test used 10 incidents that were also present in the agent’s Hindsight memory bank. For each query, the agent found the associated root cause, warned about the relevant trap action, and cited the incident. That looked like a perfect score, but the setup let the agent search information it had already been given.
“If the test data is in memory, you’re testing lookup.” That distinction matters: retrieval from stored postmortems can be useful, but it does not establish that an agent can diagnose a novel incident from symptoms. The score answered a narrower question than I intended.
#1 Best Overall
How I rebuilt the evaluation
- Keep evaluation incidents out of memory. From the 114 OpenSRE incidents in the set I used, I retained 104 in a fresh memory bank and reserved 10 as held-out cases.
- Write queries from symptoms. I made symptom-only queries for the held-out incidents instead of copying root-cause language from their postmortems.
- Run two conditions. I asked each query with memory enabled, then ran a baseline using the same model and prompt with the memory block removed. Keeping those conditions aligned makes the memory/no-memory comparison more informative.
- Grade against a defined reference. I checked classifications against each incident’s
true_categoryfield. I counted an answer as a partial match when it named a plausible cause in the right area but got the mechanism or trigger wrong. - Keep the outputs. I saved the responses in
eval_holdout_results.json, so the reported scores could be checked against the evaluation artifacts.
What the two conditions produced
| Condition | Reported result | How to read it |
|---|---|---|
| Memory enabled | 9 of 10 correct root-cause classifications; one run was blocked by Groq’s daily rate limit and counted as a miss. | A strong result on these held-out cases, with the rate-limited run included as a failure. |
| Memory removed | 0 of 10 fully correct; four partial matches and six hallucinated responses. | The baseline produced no fully correct classifications under this grading, though four answers were plausible at a broad level. |
These are my results from one small evaluation, not an independent benchmark. A demo query about checkout-service 500 errors after a deployment illustrates the difference without proving it: the no-memory model invented a NullPointerException, log counts from a kubectl command it had not run, and a Helm revision that did not exist. The memory-backed answer suggested a dependency-capacity problem and cautioned against rollback based on similar incidents, but also included irrelevant network and systemd checks.
The confidence label in that response was extracted from its text with a regular expression. It was not a calibrated probability and should not be read as one.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What this evaluation doesn’t show
The gap between 9/10 and 0/10 fully correct is an outcome in this particular setup. It does not establish how stable the result would be across repeated runs, or which part of the memory system produced the difference.
Quick Recap
Best Value
Rank #4
- Only 10 held-out cases: a small sample can give an unstable estimate of performance.
- One grader: I graded the answers myself, so there was no independent check on category decisions.
- Category-level grading: matching
true_categorydoes not mean the agent reconstructed the exact event or mechanism. - Subjective partial credit: whether a response is a plausible cause in the right area can involve judgment.
- Related cases: held-out incidents came from the same dataset and vendor set as the retained incidents; they may not be very different in practice.
- One run per query: there is no measure of run-to-run variation.
- One rate-limited run: counting the blocked memory-backed query as a miss is transparent, but it also means that score includes a service failure rather than only a diagnostic answer.
- No component ablation: the comparison does not separate reflection, recall, trap boosting, and signature enrichment.
A practical checklist for testing an incident-response agent
- Reserve incidents before seeding memory; do not let evaluation cases enter the agent’s memory.
- Use a baseline with the memory content removed while keeping the model and other prompt conditions consistent.
- Write questions from symptoms, avoiding root-cause wording that gives away the answer.
- Set grading rules before reviewing outputs, and distinguish broad category matches from correct mechanisms and triggers.
- Report failed, blocked, and rate-limited runs rather than quietly excluding them.
- Save raw outputs alongside scores so readers can inspect what was actually graded.
- For broader claims, use more cases, repeat runs, get an independent grader, and select a held-out set deliberately different from the retained incidents.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

