What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can build a small vector database in Python to learn how records, distance metrics, nearest-neighbor search, persistence, and filtering fit together. The working version in this guide uses exact cosine-distance search over in-memory records and saves them as JSON. It is an educational prototype—not a production database—and it does not implement HNSW or IVFFlat indexing.

That distinction matters: an exact scan is a useful correctness baseline, while approximate indexes trade some accuracy for search efficiency. The ten steps take you from a minimal working store to understanding what you would need to add for larger workloads.

1. Choose the scope before writing code

“From scratch” can mean different things. Here, it means implementing the core storage and search behavior yourself, using Python’s standard library—not building on PostgreSQL, pgvector, or another database engine. The project is an in-memory prototype with JSON persistence, one similarity metric, metadata filters, and exact search.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is suited to learning with small datasets. It does not provide concurrent transactions, crash-safe recovery, an approximate-neighbor index, replication, or sharding. Pick a fixed vector dimension for your records; the example below sets it when creating the database.

2. Define what a vector record contains

Each record needs a stable ID and a vector of the same known dimension. Optional metadata lets queries restrict results—for example, to records in a particular category. The database must reject vectors with the wrong number of components, non-numeric values, or non-finite numbers; otherwise, errors can surface later as misleading search results.

The example stores records in a Python dictionary keyed by ID. This makes lookup, replacement, and deletion by ID straightforward, but it does not create a similarity index: finding nearest neighbors still requires examining every eligible record.

3. Implement and understand cosine distance

A similarity search needs a defined metric. This tutorial uses cosine distance, calculated as 1 - cosine_similarity. Smaller distances rank as closer matches. Cosine distance is not the same value as cosine similarity: a similarity of 1 corresponds to a distance of 0.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The cosine formula divides the dot product by the product of the vectors’ magnitudes. A zero vector has no defined cosine direction, so the implementation rejects it rather than silently assigning it a score. This is one metric choice, not a universal recommendation: the right metric depends on how the embeddings were produced and how they should be compared.

4. Build the exact top-k search baseline

Exact search calculates the distance from a query vector to every record that passes the metadata filter, sorts the results, and returns the first k. That exhaustive scan gives the exact nearest neighbors for the chosen metric and current dataset. Its work grows with the number of records, which is why it becomes costly as the collection grows.

Save the following as vector_db.py. It requires Python 3 and uses only the standard library. The ordering is deterministic: equal distances are ordered by ID.

import json
import math
from pathlib import Path


class VectorDB:
    def __init__(self, dimension):
        if not isinstance(dimension, int) or dimension <= 0:
            raise ValueError("dimension must be a positive integer")
        self.dimension = dimension
        self.records = {}

    def _vector(self, values):
        vector = [float(value) for value in values]
        if len(vector) != self.dimension:
            raise ValueError(f"expected {self.dimension} dimensions")
        if not all(math.isfinite(value) for value in vector):
            raise ValueError("vector values must be finite")
        if sum(value * value for value in vector) == 0:
            raise ValueError("cosine distance is undefined for a zero vector")
        return vector

    @staticmethod
    def _cosine_distance(left, right):
        dot = sum(a * b for a, b in zip(left, right))
        left_norm = math.sqrt(sum(a * a for a in left))
        right_norm = math.sqrt(sum(b * b for b in right))
        return 1.0 - dot / (left_norm * right_norm)

    def upsert(self, record_id, vector, metadata=None):
        if not isinstance(record_id, str) or not record_id:
            raise ValueError("record_id must be a non-empty string")
        self.records[record_id] = {
            "vector": self._vector(vector),
            "metadata": metadata or {},
        }

    def delete(self, record_id):
        return self.records.pop(record_id, None) is not None

    def search(self, query, k=5, where=None):
        query = self._vector(query)
        if not isinstance(k, int) or k <= 0:
            raise ValueError("k must be a positive integer")
        where = where or {}
        results = []
        for record_id, record in self.records.items():
            if not all(record["metadata"].get(key) == value
                       for key, value in where.items()):
                continue
            distance = self._cosine_distance(query, record["vector"])
            results.append({"id": record_id, "distance": distance,
                            "metadata": record["metadata"]})
        results.sort(key=lambda item: (item["distance"], item["id"]))
        return results[:k]

    def save(self, filename):
        payload = {"dimension": self.dimension, "records": self.records}
        Path(filename).write_text(json.dumps(payload), encoding="utf-8")

    @classmethod
    def load(cls, filename):
        payload = json.loads(Path(filename).read_text(encoding="utf-8"))
        db = cls(payload["dimension"])
        for record_id, record in payload["records"].items():
            db.upsert(record_id, record["vector"], record["metadata"])
        return db


if __name__ == "__main__":
    db = VectorDB(dimension=3)
    db.upsert("doc-1", [1, 0, 0], {"topic": "science"})
    db.upsert("doc-2", [0.9, 0.1, 0], {"topic": "science"})
    db.upsert("doc-3", [0, 1, 0], {"topic": "travel"})

    print(db.search([1, 0, 0], k=2))
    print(db.search([1, 0, 0], k=5, where={"topic": "science"}))
    db.save("vectors.json")

Run it with python vector_db.py. The first search should return doc-1 and then doc-2; the second returns only records whose topic metadata equals science. The stored distance is the value used for ranking, not a similarity score.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Add a simple ID index, and know what it does

The records dictionary is already an index from ID to record: it supports direct ID lookup and replacement without scanning a list. It does not speed up nearest-neighbor search because a query vector has no direct dictionary key. Keeping those two jobs distinct prevents a common misconception: an ID index is not a vector index.

For a small collection, the exact scan is often the simplest option to reason about. It also gives you a result set to compare against when you later try an approximate method.

6. Understand approximate indexes and their trade-offs

Approximate nearest-neighbor indexes avoid comparing a query with every vector. Two documented options in pgvector illustrate different designs. HNSW arranges vectors in a multilayer graph; IVFFlat partitions them into inverted lists. These are pgvector characteristics, not guarantees that every implementation or workload will behave the same way.

Index How it organizes vectors Documented trade-offs and setup
Exact scan Compares the query with every eligible record. Exact nearest neighbors for the chosen metric; the work grows with the collection because every eligible row is scanned.
HNSW A multilayer graph. pgvector describes generally better speed/recall behavior than IVFFlat, with slower index builds and greater memory use. It has no training step and can be created on an empty table. Its m and ef_construction settings affect graph construction effort; greater construction effort can improve recall while increasing build time and slowing inserts.
IVFFlat Vectors are assigned to inverted lists. pgvector advises creating the index after loading data. Search examines selected lists rather than necessarily scanning every vector, so results can differ from exact search.

Neither approximate index is a universal winner. The outcome depends on the data, query workload, settings, and resource limits. This Python project intentionally stops before implementing either data structure; a convincing HNSW or IVFFlat implementation is a substantially larger project than a short addition to the exact scan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In pgvector, the distance operator and index operator class must match. Its documented operators include <-> for L2 distance, <#> for negative inner product, <=> for cosine distance, and <+> for L1 distance. For binary vectors, it also documents <~> for Hamming distance and <%> for Jaccard distance. These PostgreSQL operators are useful reference points, not syntax used by the Python code above.

7. Persist records and handle mutations

The example’s save method writes a JSON snapshot; load restores records and revalidates vectors. upsert replaces a record with the same ID, and delete removes it. This is enough to demonstrate persistence and mutations, but a single snapshot write is not a transaction log or crash-safe recovery system. Concurrent writers can also conflict because the example has no locking or transaction isolation.

For a larger store, persistence design needs to address atomic writes, recovery after interruption, schema changes, and consistency between stored vectors and indexes. If an index is derived from records, updates and deletes must maintain it or trigger a rebuild; stale index entries can return wrong or missing results.

8. Add filters and define the query contract

The search method accepts k and an optional equality filter named where. It validates the query dimension, rejects invalid k values, applies metadata filtering, then ranks eligible records. Its filter supports only exact equality on top-level metadata keys; range conditions, nested fields, and a richer query language are out of scope.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In this exact implementation, filtering happens before distances are calculated, so the method returns up to k matching records if enough exist. Approximate systems can behave differently: if they scan a limited candidate set and apply a selective filter afterward, they may return fewer than the requested number. Supabase documents iterative scans for pgvector 0.8.0 and later as one way to search further for enough filtered results; actual behavior depends on configuration and limits.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Benchmark an approximate method against the exact baseline

Do not call an index faster or accurate enough without measuring it on the workload that matters. Use exact search as the ground truth, run the same queries through the approximate index, and record the difference along with resource costs.

  • Recall at k: for each query, divide the number of approximate results that also appear in the exact top-k by k, then average across queries. If fewer than k matching records exist, use the number of available exact results as the denominator.
  • Query latency: measure elapsed time across a representative query set, not just one favorable query.
  • Build and update cost: record index build time and the effect of inserts, replacements, and deletes.
  • Memory and disk use: measure the resources required by the index and stored data.
  • Filter behavior: test both common and selective filters, checking whether the requested number of results is returned.

Report the dataset, vector dimension, hardware, index settings, and query mix alongside the measurements. There is no universal speedup or recall figure established for all vector databases.

10. Know what a production system still needs

Nearest-neighbor retrieval is only one database responsibility. Production choices depend on requirements that this prototype leaves open:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • More data or lower memory use: pgvector documents half-precision vectors with halfvec and binary quantization with reranking. Reduced representations can change precision and need evaluation against the original vectors.
  • Keyword plus semantic retrieval: hybrid full-text and vector search can combine lexical matching with embedding-based ranking.
  • Operational safety: concurrency, crash recovery, replication, and scaling require deliberate designs beyond a JSON snapshot. A 2026 research system, PostgreSQL-V 2.0, addresses concurrency, crash recovery, and physical replication in its PostgreSQL integration; its results are specific to that prototype and its benchmarks.
  • PostgreSQL implementation practices: pgvector recommends bulk loading with COPY, creating indexes after initial loading where appropriate, inspecting plans with EXPLAIN (ANALYZE, BUFFERS), and using concurrent index creation in production to avoid blocking writes. Those are PostgreSQL-specific practices, not requirements for this Python prototype.

For a managed PostgreSQL example, Google Cloud SQL documents storing, querying, and indexing embeddings with pgvector, including HNSW index creation. That is one provider’s implementation path, not a prerequisite for learning vector search.

What this project teaches—and what it does not

You now have a small store that validates fixed-dimension vectors, ranks exact cosine-distance results, filters by metadata equality, supports replacement and deletion, and can save and reload a snapshot. The exact scan is the key reference point: any approximate index you add should be evaluated against it for recall as well as latency and resource use. Turning this prototype into a production database would require substantially more work on indexing, durability, concurrency, and operations.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.