The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Use Amazon Kinesis Data Streams as the event backbone for a stateful AI agent—not as the agent’s memory database. Producers write conversation turns, tool results, preference changes, and domain events to a stream. Consumers process those events into durable profiles, summaries, vector indexes, or knowledge graphs. When the agent runs, it retrieves only the authorized information relevant to that request.
What Kinesis does—and what it does not do
Kinesis Data Streams transports and retains records so they can be consumed and replayed. AWS describes a stream as a set of shards and a data record as the unit stored in a stream. A record has a sequence number, partition key, and data blob; AWS documentation published in 2026 states that a blob can be up to 1 MB.
That makes Kinesis useful for capturing a history of events, but it does not interpret those events as memories. A model cannot efficiently or safely use a raw stream as its conversational context. A separate consumer must turn events into queryable state, and the agent must retrieve a small, relevant subset of that state when needed.
How to build agent memory from a Kinesis stream
- Define events. Decide which changes matter: user messages, tool outputs, explicit preference updates, task outcomes, or business-system changes. Give each event a stable identifier and a clear type. A practical event shape might include
event_id,tenant_id,user_id,event_type,occurred_at, and a type-specificpayload. Avoid treating every transient model thought or full prompt as a permanent memory by default. - Write events to the stream. Producers can use the Kinesis APIs, including
PutRecordandPutRecords, the Kinesis Producer Library, or Kinesis Agent for file-based ingestion. Choose a partition key that reflects the ordering you need. For per-user ordering, a stable key such astenant_id:user_idis a useful starting point. - Process records with an appropriate consumer. Use Lambda for managed, record-oriented handlers; Kinesis Client Library (KCL) when you need a custom long-running consumer and control over processing; or Managed Service for Apache Flink for stateful or windowed stream processing. Firehose can deliver stream data to downstream destinations. These options are not interchangeable: choose based on processing state, control, latency, and operational needs.
- Build durable projections. Have consumers update the stores that answer agent queries: a profile store for explicit preferences, a summary store for compact conversation history, a vector index for semantic retrieval, or a context or knowledge graph for relationships and structured facts. Store version or sequence metadata with projections so that replaying an older event cannot overwrite newer state.
- Retrieve context at invocation time. Before the model sees anything, fetch the profile facts, recent summaries, and task-relevant events needed for the request. Enforce tenant isolation and authorization during retrieval, not as a prompt-only instruction. Keep the assembled context compact rather than sending the entire stream or all stored history.
- Plan for replay and recovery. Keep raw events replayable for the retention period your system requires, and document how to rebuild each projection. Make updates idempotent so retrying a record does not create duplicate memories or corrupt state. Define how corrections, deletions, and expired data propagate from the event history to every projection.
Choose the consumer model for the work
| Option | Best fit | Trade-off |
|---|---|---|
| Lambda | Managed, record-oriented event handling with less infrastructure to operate | Less control than a custom long-running consumer; design the handler around retries and idempotent writes. |
| KCL consumer | A custom consumer that needs control over processing and checkpointing | You operate the consumer service and its deployment, scaling, and recovery behavior. |
| Managed Service for Apache Flink | Stateful transformations, event windows, or continuous stream processing | More processing capability, with corresponding application and operational complexity. |
The right choice depends on the projection workload. A simple preference update may suit a Lambda handler; a stream that calculates rolling state over event windows may call for Flink. KCL is appropriate when managed record handling is not enough and you need a custom consumer. Test the actual processing and recovery path, not only the happy-path event flow.
#1 Best Overall
Choose partitioning and read capacity deliberately
Partition keys and ordering
A partition key determines how records are assigned across the stream’s shards. A stable per-user key can help preserve the ordering needed to update that user’s state. But sending too much traffic under one key can concentrate work on a hot shard and limit parallelism. The shard layout and resharding plan therefore affect both throughput and how many consumers can work in parallel.
Shared reads or enhanced fan-out
With shared consumption, consumers share shard read capacity. Enhanced fan-out gives each registered consumer dedicated read throughput: AWS documentation published in 2026 states 2 MB per second per shard per enhanced-fan-out consumer and describes delivery typically 70 milliseconds after stream arrival. AWS recommends considering it when multiple consumers need parallel reads or when using SubscribeToShard for low-latency delivery. It is a capacity and latency choice, not a prerequisite for agent memory.
Rank #2
On-demand or provisioned capacity
On-demand mode reduces the need to plan shard capacity in advance, while provisioned capacity lets you plan explicitly around shard allocation and expected load. AWS documentation published in 2026 lists on-demand write capacity as starting at 4 MB per second and 4,000 records per second, with default scaling up to 200 MB per second and 200,000 records per second. Treat those as documented service capacity figures, not a guarantee that a particular workload will achieve a specific end-to-end rate; actual needs depend on record size, traffic shape, partition distribution, and consumers.
Pick the memory store to match the question
| Projection | Use it for | Important limitation |
|---|---|---|
| Structured profile | Explicit preferences and stable facts that should be read deterministically | Requires a schema and rules for updates, conflicts, and deletion. |
| Summary store | Compact continuity across longer conversations or completed tasks | A summary is a derived view; retain source events if you need to audit or rebuild it. |
| Vector index | Semantic search over relevant past events or documents | Similarity is not authorization or truth; validate access and treat retrieved text as evidence to assess, not unquestionable fact. |
| Context or knowledge graph | Relationships and linked facts that benefit from structured traversal | Requires explicit modeling and maintenance of entities and relationships. |
Many systems combine these projections. A vector search can surface a relevant past discussion, while a structured profile supplies a deterministic preference or permission check. The stream remains the event history; projections are purpose-built ways to answer different retrieval questions.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteMake retries, checkpoints, and replay safe
Stream processing can retry work, and a consumer may resume from a checkpoint after interruption. Treat delivery as something that can be repeated: use stable event IDs, idempotent projection writes, and version checks. A projection update should reject or safely ignore an event older than the version already applied.
- Checkpoint only after durable work: a checkpoint should represent progress the consumer can recover from without losing an event or advancing past a failed projection update.
- Handle failed records explicitly: record failures for investigation and retry or route them through a defined recovery path. Do not silently mark an event processed if its memory projection was not updated.
- Test rebuilds: replay a representative event range into a clean projection and verify that the result matches expected state. Include duplicate, late, corrected, and deleted events in the test cases.
- Monitor freshness and lag: track iterator age, read and write throttling, checkpoint lag, duplicate handling, failed records, and projection freshness. A healthy stream does not guarantee that the memory store is current.
For file-based ingestion, AWS Kinesis Agent documentation describes checkpointing, retries, and CloudWatch metrics. For any ingestion path, identify which component owns retry behavior and how operators can distinguish a delayed projection from a lost or rejected event.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Protect privacy and tenant boundaries
Conversation events can contain personal, confidential, or regulated information. Decide what may enter the stream, how long raw records and derived projections remain available, and how a user’s deletion request reaches every store. Restrict access to the stream and projections according to tenant and service boundaries. At retrieval time, apply authorization before constructing the model context; filtering only after the model receives content is too late.
Keep the event record useful for replay without unnecessarily retaining full prompts or sensitive tool output. Where a projection derives a fact from an event, preserve enough provenance to find the source and correct or remove the derived fact when required.
Best Value
Operational checklist
- Partition keys reflect the ordering requirement without concentrating all traffic on a hot key.
- Consumers have explicit checkpoint, retry, duplicate, and failed-record behavior.
- Projection writes are idempotent and protected by version or sequence checks.
- Replay and projection rebuild procedures have been tested.
- Dashboards cover throttling, consumer lag, failures, and projection freshness.
- Retention, deletion, access control, and tenant isolation cover raw events and every derived store.
- Agent retrieval returns only the authorized facts relevant to the current request.
What to evaluate before committing
Kinesis is a good fit when the architecture benefits from a durable event stream that multiple consumers can process and replay. It does not by itself solve semantic retrieval, profile conflict resolution, memory summarization, or authorization. Capacity mode, retention needs, consumer count, and downstream stores all affect the design and its cost. AWS has not published a title-specific head-to-head benchmark establishing that Kinesis is cheaper or faster than every competing broker, so compare candidate architectures using the event sizes, throughput, latency targets, replay needs, and retrieval workload you actually expect.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

