iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A system can remember an account’s entire history without sending that history to the model on every request. Building Waada, a sales-handover application, taught me to treat persistent memory and prompt context as separate design problems: memory preserves evidence; retrieval selects the evidence a particular task needs; and a context budget limits what reaches the model. This is an account of one implementation, not proof that memory-aware systems always beat simpler approaches.
Memory is not the same as prompt context
Waada is designed to help hand sales accounts from one person to another. Its inputs include email, Slack conversations, call and meeting transcripts, audio, and CRM information. Those sources can contain a long history, but adding all of it to every model call would be wasteful and could make the answer less useful. A model asked about an outstanding promise needs different evidence from one asked what changed recently.
The distinction is between durable storage and a temporary reasoning window. As I put it in the original account, “The memory store should remain the source of historical evidence. The prompt is a temporary reasoning window.” Keeping more history available does not mean including more history in each prompt.
In Waada, Hindsight is the separate memory and retrieval layer. The application asks it for evidence, then the LLM reasons over what comes back. That describes this implementation; it is not a claim that every Hindsight deployment works identically or that Hindsight is the only way to build such a system.
#1 Best Overall
Normalize source material before retrieving it
Before retrieval can select evidence consistently, Waada converts source-specific interactions into a canonical structure. The described fields include the account, source ID, type, date, title, participants, content, and source metadata. The reason is practical: if email, CRM records, and transcript chunks are represented in incompatible ways, the application has a harder time selecting and bounding relevant evidence.
Dates and source context must survive this transformation. A passage stripped of its date may appear current when it is old; one stripped of its source can be harder to interpret or verify. A compact prompt is not necessarily a reliable prompt if it loses the context needed to judge what a claim means.
Start retrieval with the question
Waada uses distinct retrieval intents rather than one generic request for “account context.” Different questions imply different evidence searches:
Recommended Free Tools
Rank #2
- Open commitments: “Which promises are still open?” searches for promise-related evidence through a commitment ledger.
- Sensitive subjects and objections: “What should the new owner know before reopening a difficult topic?” calls for objections, sensitive subjects, and agreements the customer has accepted.
- Recent developments: “What changed since July?” requires evidence selected for changes over time, with dates available to judge recency.
- Stakeholders or a direct question: These call for their own retrieval paths, rather than assuming the same evidence is useful for every task.
The design implication is that retrieval should be shaped by the work the model is about to do. A single broad account summary can miss a specific promise, bury a sensitive objection, or emphasize history that does not answer the question.
Retrieve first, then enforce the context budget
Retrieval produces candidates, not a finished prompt. In the described flow, the application deduplicates the returned evidence, chunks it, selects useful material, and caps what it sends to the LLM. That order matters: a budget cannot guide selection if the system has not first determined which evidence could answer the task.
The author reports that Waada’s LLM layer uses a 5,000-token input budget and describes this conservatively as approximately 12,500 characters. The character figure is not a precise tokenizer measurement, and both numbers describe this project’s configuration, not a universal model or provider limit. Token counts vary with the content and tokenizer, so a character approximation should not be treated as an exact conversion.
“Provider limits are part of application architecture,” I wrote. A bounded prompt is not merely a way to avoid a rejected request; it makes the application choose what matters and what can safely be left out. The author’s stated preference was: “I would rather drop low-value context deliberately than let the prompt grow until the provider rejects it.”
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Validate generated output before using it as state
Retrieval and generation do not guarantee that the model will return usable structured data. Waada validates structured extraction with Zod. When the output is invalid, the described implementation can provide repair guidance and try JSON parsing; if the attempts fail, the result can be null rather than being accepted as application state.
This makes validation part of the control flow, not an optional cleanup step. A malformed response should be treated as a recoverable failure, not silently trusted because it came from a model. Returning no result is safer than converting invalid data into a false commitment or account fact.
Rank #4
What the comparison did—and did not—show
The author describes evaluating three approaches: CRM-only evidence, raw-summary-only context, and memory-aware retrieval. This was not a benchmark of commercial CRM products, and the account does not establish a universal accuracy advantage for memory-aware systems. The reported live evaluation was mixed: some retrieval behavior worked as intended, while structured-output variability, rate-limit pressure, prompt-size problems, and intermittent Hindsight failures also appeared. In one run, the summary-only baseline scored higher on the checks.
That outcome is important because it leaves room for a simpler baseline to win on a particular run. A useful evaluation should record degraded conditions and failures as well as successful retrieval, and compare the actual task outcomes rather than assuming that a more elaborate memory layer is automatically better.
Rate limits and service failures belong in the design
The implementation ran operations sequentially in at least part of the system to fit a provider’s shared rate window. The author presents this as predictable control flow under rate limits, not as a general performance optimization. A design that avoids bursts may be more manageable under a specific limit, but sequential work can also affect how long a workflow takes.
Best Value
Intermittent memory-service failures and prompt-size problems were also part of the reported evaluation. They are reminders to define what the application does when evidence cannot be retrieved or a provider rejects a request. The account does not specify a universal recovery policy; the relevant principle is to keep such failures visible and prevent them from being mistaken for successful, evidence-backed answers.
The architectural lesson
Persistent memory answers what historical evidence can be retained. Retrieval answers which evidence is relevant to this task. Context budgeting answers how much of that evidence should reach the model. Validation and failure handling determine whether the generated result is safe for the application to use.
In the author’s words, “In a memory-aware application, context isn’t just input. Context is architecture.” The value of the design is not the sheer size of its memory, but the discipline to select evidence, preserve provenance and dates, respect operational limits, and reject output that does not meet the application’s requirements.
Read the original DEV Community account of building Waada.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

