Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Sometimes—but “millions of vectors on 2GB” is not a dependable promise without testing your exact data, extension, and VPS. SQLite stores its database in files, but the application, SQLite cache, vector index, active queries, and operating-system file cache still compete for memory. Start by calculating the vector payload, choose a search method that fits your recall and latency needs, then benchmark under the VPS’s actual memory and storage limits.

What 2GB of RAM does—and does not—tell you

SQLite’s official overview describes it as an in-process database engine: it reads and writes ordinary files rather than running as a separate database server. That can make it practical on a small machine, but it does not make vector search memory-free. The application and SQLite extension use memory, and the operating system may use spare RAM to cache database pages. A file-backed database can be larger than RAM; whether queries remain acceptably fast depends partly on how much data must be read from storage.

Think of the VPS’s memory as a whole-system budget, not a quota reserved for vectors. The application, SQLite, extension and index, concurrent queries, and operating system all need room. A database-size ceiling is not a RAM guarantee, and a vector count alone cannot establish whether a workload will fit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Estimate the vector payload before choosing an index

For uncompressed float32 vectors, estimate the payload as:

vector count × dimensions × 4 bytes

That is just the raw vector values. It excludes SQLite row and page overhead, metadata, any separate index structures, temporary allocations during queries or index building, and the application’s own memory.

Example payload Calculation What the estimate means
1 million, 768-dimensional float32 vectors 1,000,000 × 768 × 4 bytes About 3.07 GB of raw values, before SQLite row/page overhead. This is the SQLite-Vector documentation’s estimate.
1 million, 768-dimensional int8 vectors 1,000,000 × 768 × 1 byte About 768 MB of raw values, calculated from one byte per element; it is not a measured database or resident-memory total.
1 million, 384-dimensional float32 vectors 1,000,000 × 384 × 4 bytes About 1.536 GB of raw values, calculated before database and runtime overhead.

The calculations use decimal units for readability. Actual on-disk and runtime memory use depends on the representation, extension, index, SQLite settings, and workload. In particular, a 3.07 GB raw payload already exceeds a 2GB machine’s physical memory before overhead, while a smaller payload still does not prove the full service will fit.

Choose a search strategy for the workload

Exact search checks candidates directly; approximate-nearest-neighbor (ANN) methods trade some search behavior for an index intended to narrow the candidates. Quantized representations reduce the number of bytes used to represent vectors, but may reduce recall. Index construction can also require time, memory, or training work. Compare approaches against the same embeddings and target recall rather than assuming one vector-count threshold determines the answer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option What the documentation describes Trade-offs to assess
SQLite-Vector Ordinary SQLite table BLOBs; documented float32, float16, bfloat16, int8, uint8, and 1-bit formats, plus exact scans, quantized scans, and a bounded streamed mode. The project says its approach does not require pre-indexing. Check the package version and actual defaults. Compare recall, scan latency, peak process and system memory, and storage reads for your vectors. Its published benchmark is not a measurement of your VPS.
sqlite-vec A small, portable C extension with vec0 virtual tables and documented float, int8, and binary vectors, along with metadata, auxiliary columns, and partition keys. The repository marks the extension pre-v1 and warns that breaking changes may occur. Verify compatibility with the SQLite build in your application and test the update path you need.
SQLite vec1 Official SQLite extension documentation describes approximate search using IVFADC with OPQ for L2 and cosine distances. Its reference also describes a flat index storing vectors in packed SQL BLOBs and a bucket-probe query parameter. The overview lists incomplete areas and roadmap items. Confirm the features you require are implemented, then measure index build time, query behavior, memory, and recall for your configuration.
sqlite-vss The repository documents Faiss-based indexes and says IVF requires training. The project says it is no longer actively developed and points to sqlite-vec as the maintainer’s focus. Its News Category example reports 45 minutes to train an IVF4096 configuration and 4.5 minutes to insert; those figures describe that example, not a prediction for your dataset.

For any option, measure exact versus approximate search, recall at your chosen k, warm and cold query latency, index build or training time, incremental write and update behavior, on-disk footprint, peak process and system memory, storage bandwidth, extension maintenance status, and compatibility with the application’s SQLite build.

What the published bounded-memory benchmark actually shows

SQLite-Vector’s project-reported benchmark covers one million vectors with an INT8 index on an Apple M5 Pro. The project reports the following results; the pages do not give a benchmark publication date, and the figures should not be read as a VPS test.

SQLite-Vector mode Peak memory Query time Recall
Preloaded 740 MB 37.6 ms per query 99.5% Recall@20
Streamed 30 MB 114.4 ms per query 99.5% Recall@20

The project characterizes streamed mode as a 25-times memory reduction for roughly three-times query latency, with the same reported recall. Its benchmark page says timings are the best of 20 runs with the database file in the operating-system page cache. It also explains that when the index does not fit in RAM, storage read time must be added according to index size and storage bandwidth.

Do not treat the reported 30 MB as the total physical-RAM requirement for a running service. The project notes that resident memory depends on file-backed versus in-memory operation, SQLite cache settings, preload mode, page-cache behavior, and the host allocator. The benchmark does not establish performance or out-of-memory safety for a particular $5 VPS.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quantization can save space, but verify recall

SQLite-Vector’s documentation gives approximate scan-representation storage ratios of 13% of the raw payload at 4-bit, 10% at 3-bit, and 7% at 2-bit. These are implementation documentation figures, not guarantees for every dataset or version. Lower storage use does not by itself establish the recall, latency, or total resident-memory behavior you will get.

In the project’s real-dataset example with 10,000 vectors, the documented Recall@10 is 0.948 at TurboQuant 4-bit, 0.868 at 3-bit, and 0.596 at 2-bit. The project recommends starting with 4-bit when recall matters and validating 2-bit against the target embeddings. Treat these results as that documented example, not as a forecast for your own vectors.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Benchmark the exact VPS before deployment

A useful test reproduces the intended deployment rather than measuring an isolated query on a larger development machine. Record the VPS model and price, available RAM, storage type and performance, operating system, application and SQLite build, extension name and exact version, vector count and dimensions, embedding source, distance metric, and any SQLite cache or preload settings. State the target recall and whether searches run concurrently.

  1. Build representative data. Use the intended embedding format, dimensions, metadata, and write pattern. Include the full expected dataset if feasible, or test staged sizes and do not extrapolate linearly without checking.
  2. Create an exact baseline. For the same query set and metric, calculate exact nearest neighbors on a machine or test configuration capable of producing a reference. Compare the candidate method’s results at the same k to calculate recall.
  3. Separate warm and cold runs. Report whether the database pages were already in the operating-system cache. Cold reads on an index that does not fit in memory can be limited by storage bandwidth; a warm-cache result does not describe that case.
  4. Measure the whole service. Track peak process resident memory and whole-system memory during startup, index construction, single queries, and expected concurrent queries. Include the application and operating system rather than reporting only an extension’s allocation.
  5. Measure performance and operations. Record query latency distributions, recall at the chosen k, build or training duration, on-disk size, and the cost of inserts or updates. Repeat under expected concurrency and with the cache condition labeled.
  6. Test the deployment limit. Run with the actual VPS memory limit and representative data, leaving room for normal application and operating-system use. Watch for memory pressure, swap activity, or an out-of-memory termination before declaring the configuration safe.

Publish the configuration alongside results so that a figure remains interpretable: hardware and storage, software versions, dataset shape, metric, recall target, concurrency, cache condition, and how memory was measured. Without those details, a claim such as “millions fit” is not useful for sizing another machine.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the design needs to change

If the measured configuration misses its recall target, exceeds the service’s memory budget, or has unacceptable cold-query latency, choose the adjustment that addresses the measured bottleneck:

  • Memory pressure: test a more compact representation or streamed scanning, then remeasure total service memory and recall.
  • Recall loss: use a less aggressive quantization level or another supported search method, and compare against the exact baseline.
  • Slow queries: determine whether time is spent scanning, reading uncached pages, or using an unsuitable index before increasing preload or changing methods.
  • Expensive index construction or updates: measure build/training and incremental-write behavior on the actual data; a fast query result alone does not make the operational design suitable.
  • Budget still fails: reduce dimensions or dataset volume, or select storage and search infrastructure designed for the measured workload. That is a design decision to validate, not a result established by the cited benchmark.

There is no supported universal vector-count cutoff for a 2GB SQLite deployment. The practical limit is the largest configuration that meets your measured memory, recall, latency, and update requirements on the exact target machine.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.