Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a small, read-only dataset that can be regenerated, a continuously running database may be unnecessary. In the third installment of his AWS serverless migration account, Dmitriy Trunov describes replacing that part of an agentic retrieval-augmented generation (RAG) application with a SQLite/FTS5 artifact in S3, while moving mutable conversation, feedback, and spending data to DynamoDB. Rebuilding the corpus also exposed identifier collisions and oversized chunks. Crucially, the migration’s end-to-end answer-quality checks had not yet run, so whether answer quality survived remained unknown.

Why remove a database from a read-only workload?

The useful distinction is not simply SQL versus NoSQL. It is whether the data must be updated during normal operation, how much of it exists, and whether it can be regenerated reliably.

In the preceding installment, Trunov described a relational database holding 278 project records and associated library rows: 88 KB in total. The data was rebuilt from scratch and read-only during queries. For this workload, his replacement was a generated projects.sqlite artifact containing tables, the corpus, and SQLite FTS5 search. The pipeline published the artifact to S3, and Lambda loaded it into /tmp. Conversation history, feedback, and spend tracking—data that changes—went to DynamoDB instead. See Trunov’s second installment.

This is a case-specific architecture, not a rule that small datasets should always use SQLite. A generated file can suit a read-heavy workload when the application can rebuild and publish it, its queries fit the embedded database’s capabilities, and the Lambda runtime can load and use it within its operational limits. The design also separates deployment and refresh of the read-only artifact from writes to mutable state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Questions to answer before choosing the same pattern

  • Can the data be regenerated? If rebuilding is unreliable or destructive, the artifact is not a safe substitute for a durable source of truth.
  • Is query behavior a fit? Trunov used SQLite FTS5 for keyword search alongside the corpus. Workloads needing different database features or write patterns may not fit.
  • What is the idle cost and resume cost? Compare the cost of a continuously available database with the storage and loading path for a generated artifact; include the delay and operational work of refreshing it.
  • What limits apply at runtime? Verify artifact size, Lambda storage and execution constraints, and the time needed to load the file for the application’s own deployment and traffic patterns.
  • Can mutable state remain separate? This design sent conversations, feedback, and spending data to DynamoDB rather than treating the generated SQLite file as a write store.

What did rebuilding the corpus uncover?

Regeneration was more than a deployment step: it gave Trunov a chance to inspect the material the old retrieval pipeline had accumulated. In his 2026 account, he reports two defects in the rebuilt corpus. These measurements are the author’s reported results, not independently verified benchmarks.

Colliding chunk identifiers

The original identifier format was {repo}::{file_path}::{section}. When headings repeated within a file, different chunks could receive the same composed ID. Trunov reports 303 colliding IDs among 24,775 chunks. Because reciprocal-rank fusion (RRF) deduplicated on that ID, a chunk sharing an identifier with another could be shadowed and fail to surface reliably. He says the fix was to add a per-file ordinal to the hash, distinguishing repeated sections in the same file.

Chunks too large for the embedding input

Trunov reports that the largest chunk before correction measured 119,786 bytes, which he estimated at roughly 30,000 tokens. His article states that Titan Text Embeddings accepts at most 8,192 input tokens. The chunk therefore exceeded the stated limit by a wide margin. Splitting at paragraph boundaries, with a hard fallback for long tables and code blocks, brought the reported maximum down to 7,998 bytes. The rebuilt corpus contained 25,482 chunks.

Bytes and tokens are not interchangeable measures: the token estimate is Trunov’s estimate, while the byte counts are his reported corpus measurements. The practical lesson is to validate chunks against the actual embedding model’s input limit rather than assuming that a parser’s section boundaries produce valid embedding inputs.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How did the mutable spending cap work?

The migration also changed spend tracking. Trunov reports porting an atomic reservation to a DynamoDB conditional write. The update adds a reservation only when the existing total leaves enough headroom. The item key includes the UTC date, making the cap daily rather than a lifetime limit and avoiding a separate reset job.

In a reported test against a DynamoDB implementation, 40 concurrent requests competed for a capacity of five reservations; exactly five were granted. That is a result for this test and implementation, not a universal concurrency guarantee. Applications should assess their own key design, conditional expression, reservation lifecycle, and handling of rejected writes.

Did the migration preserve answer quality?

That remained unmeasured in Trunov’s account. Two user-facing quality gates had not produced results:

  • Tool-routing comparison: question-by-question comparison against the OpenAI baseline had not run; the author says it depended on model access.
  • Retrieval remeasurement: hit rate and mean reciprocal rank had not been remeasured after replacing MiniLM/minsearch with Titan, S3 Vectors, and SQLite FTS5; the author says this depended on a generated ground-truth set.

Trunov separately reports narrower implementation checks: testing DynamoDB behavior, porting SQL behavior against a real 279-project artifact, and exercising keyword retrieval over the rebuilt 25,482-chunk corpus. Those checks help establish that components behaved in the tested cases. They do not establish that routing decisions or retrieved passages remained as good for users, and they do not substitute for the missing end-to-end quality gates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical quality-preservation checklist

  • Keep a representative set of questions and expected tools or evidence from the baseline.
  • Compare routing decisions question by question, recording where the new system chooses a different tool.
  • Evaluate retrieval against a labeled ground-truth set; report hit rate and mean reciprocal rank rather than relying on successful search calls alone.
  • Inspect failures as well as aggregate metrics, including repeated headings, long tables, and code-heavy files.
  • Separate component tests from user-facing evaluation, and label any unrun gate as unknown rather than implying equivalence.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What this case does—and does not—show

Trunov’s migration illustrates a useful decision: keep rebuildable, read-only data in a generated artifact when its query needs and runtime fit, and put mutable application state in a service designed for writes. It also shows that re-deriving data can audit it: rebuilding exposed identifiers that were not unique and chunks larger than the stated embedding input limit.

It does not show that this architecture will be cheaper or better for every application, nor that the RAG app preserved answer quality. The account’s quality gates were still outstanding. The author’s framing, “Price the floor, not the feature,” points to an architecture trade-off: compare the baseline cost of keeping services available with the actual storage, loading, refresh, and operational requirements of the alternative—not just a feature checklist.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.