Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
The case that memory, rather than scale alone, is the next step for large language models rests on one developer’s experiment, not on a field-wide consensus. In a data analysis report dated 23 September 2026, the author, writing under the byline uos1231234 on DEV Community, reports that a memory system recovered 128 of 128 target mappings in its final run, across a synthetic corpus of roughly 3 million token-equivalents. That is a striking result for a single system. It does not show that memory is the consensus next frontier for LLMs, and it does not show that the approach works on other models or on ordinary conversations.
What the author is arguing
The thesis is stated plainly: “In my view, the LLM’s next step should be memory — giving LLMs a human-like memory mechanism instead of only an attention mechanism.” This is the author’s opinion, and the experiment is offered as evidence for one implementation of it rather than as proof of the general claim.
The proposed design has two parts. First, prior interaction is compressed into useful state. Second, when that state is incomplete, the system retrieves the original details. The author’s model is deepseek-v4-flash, accessed through a tao-deepseek relay. Compression, archiving and recall were carried out by a system agent running the same model.
What was tested
The unit under test is the whole system, not a single recall component. It combines tiered compression, envelopes, tombstones and a recall fallback. Because these parts run together, the results cannot be credited to any one of them. The author’s own analysis makes this distinction: the final recovered needles are attributed to keys being preserved through the compression path, not to a recall tool call at answer time.
#1 Best Overall
| Element | Reported value |
|---|---|
| Corpus size | Approximately 3,007,411 token-equivalents across 64 blocks (a token-equivalent estimate, not a provider-metered count) |
| Block size | About 118.5K characters per block |
| Distractors | 384 blocks drawn from the same distribution |
| Targets | 128 golden service-to-key mappings |
| Model | deepseek-v4-flash via tao-deepseek relay |
| Source | “MRCR-3M Constrained Recall Experiment — Data Analysis Report (2026-09-23)”, uos1231234, DEV Community |
Results by run
The author ran several iterations. Scores are exact-key hits out of 128, using the key-string rule described below.
| Run | Score | Percentage | What the author reports changed |
|---|---|---|---|
| r1 | 116/128 | 90.6% | Initial baseline |
| r2 | 118/128 | 92.2% | Losses traced to envelope sections swallowed by a chunker (r1–r3) |
| r3 | 114/128 | 89.1% | Same chunker issue persisted |
| r7 | 126/128 | 98.4% | Fence, retry, tombstone and reminder changes |
| r7-clean | 126/128 | 98.4% | Clean variant of the r7 setup |
| r8 | 128/128 | 100% | Final run; the final ASK round used zero tool calls |
The early losses were, in the author’s account, a chunking problem rather than a memory failure: envelope sections were being swallowed before they reached the compressed state. The changes made in r7 raised scores to 126/128. The last two needles recovered in r8 are attributed to key preservation during compression.
Rank #2
How the score is counted
The headline percentages use a key-string rule: a target key counts as recalled if the key string appears in the model’s reply. The write-up also describes a stricter measure that requires the service name and key to appear together on the same line. The percentages above use the looser rule, so they should be read as a measure of whether the right key surfaced, not of whether it was tied to the right service in a strict format.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe corpus is synthetic and built for this system. It is not a standardized public benchmark, and the results have not been independently audited or replicated.
Bookkeeping checks the author reports
- Tombstones: 61 retained, from 51 compressions and 11 M3 batches.
- Paged reads: a 61-message mailbox was read in two pages (50 + 11).
- Targeted probe: zero limit collisions and zero violations, according to the author.
- Software test counts: 3,059 passing tests in the listed baseline. A later count lists 3,062 total tests, with one skipped, one todo and one stale environment failure. The stale failure is noted in the write-up itself and is not a memory result.
Problems the author reports
The write-up is useful because it records failures, not only successes. Four are worth knowing before applying any of this design.
A fixed 512K reminder fired early
The local token counter told the system that session context exceeded 512K tokens. Provider prompt-token readings in blocks 10, 12 and 13 were 202,650, 212,223 and 232,036 respectively, roughly 210K. The author attributes the gap to the local counter overestimating tokens on repetitive material. Any threshold built on a local estimate needs calibration against provider readings first.
An index went stale during compression
An index built on the length of the mutable history became invalid when compression rewrote that history during the ASK interval. The lookup was rescued by a fallback that scanned history for the last assistant string reply. The fallback produced the right answer, but the index design was fragile, and the fallback was doing essential work.
Free tools Windows power users keep installed
One-click scans. No signup required.
A revival test replayed an old answer
A revival test initially returned an existing answer byte for byte, so it did not test recall. The author removed the questions and answers and reran the test on a clean state.
Best Value
Prior answers contaminated later runs
Earlier answers remained in the history and leaked into later runs. The author removed this prior-answer contamination before the clean runs, a cleanup the write-up calls “debridement.”
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Recall accuracy is not the same as reasoning
A 2026 paper in the ACL Anthology makes the broader point. Retrieval-centric tests can fail to establish that a model reasons over long contexts, and they can be vulnerable to leakage, short-circuiting and setups that are easy to identify. More complete evaluations include multi-hop inference, aggregation and reasoning about absent information. The experiment described here is an exact-key retrieval test. It shows that the right keys were found in a synthetic corpus, and it does not test any of those harder abilities.
How to judge a memory design
The experiment is more useful as a checklist of questions than as a ranking. Comparing memory designs, whether this one or others, requires asking the same five things.
Recommended Free Tools
| Axis | Question to ask | What this experiment shows |
|---|---|---|
| What carries information | Does the state come from a summary, a raw archive, or both? | A compressed path carried the final keys, according to the author. |
| When recall is invoked | Is recall always used, used only when summaries are incomplete, or used after a check? | A recall fallback exists and was needed after an index failure; the final ASK round used zero tool calls. |
| Evaluation breadth | Does the test go beyond exact-string retrieval? | Exact-key retrieval only. Multi-hop, aggregation and absence reasoning were not reported. |
| Reliability controls | How are lineage, contamination and token counts handled? | Tombstones were used and contamination was removed; the local counter was found to overestimate on repetitive material. |
| Generalization | Has the result been reproduced across models, datasets and task types? | One system, one model configuration, one synthetic corpus. No independent replication is reported. |
What is and is not established
- Established within the author’s report: one memory system, using one model, recovered 128 of 128 key mappings under the key-string rule in its final run, with zero tool calls in the final ASK round.
- Not established: independent replication, performance on other models, performance on ordinary user conversations, superiority over alternative memory designs, or reasoning over long contexts.
- Not established as a field consensus: that memory is the next step for LLMs. That is the author’s thesis, supported by one experiment.
A reader who wants to test the idea should look first at the fallback and calibration failures described above. Those are where a memory design is most likely to break in practice.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

