Short answer: choose pgvector when PostgreSQL is already your system of record; Chroma for a simple developer-first local or small RAG project; LanceDB OSS for embedded or multimodal retrieval; Qdrant for a focused production vector service with demanding metadata filters; Weaviate for object-centric hybrid search; Milvus when distributed scale justifies additional infrastructure; and Vespa when vector retrieval is one part of a broader search, ranking, and serving platform.
There is no universal winner. The important decision is not usually which database has the fastest unfiltered nearest-neighbor query. It is where vectors belong in your architecture: inside an existing relational database, inside an embedded library, inside a dedicated vector-search service, or inside a full search and serving platform.
This comparison focuses on the open-source projects themselves. A managed cloud service built around an open-source core may add proprietary automation, support, authentication, backups, scaling, or other features, and its commercial terms can change independently of the project license.
What counts as a vector database?
A vector database stores numerical embeddings and makes them searchable by similarity. In practice, modern systems also need metadata filters, persistence, updates, access controls, lexical search, hybrid retrieval, reranking, backups, and a workable deployment model.
#1 Best Overall
- Get NVMe solid state performance with up to 1050MB/s read and 1000MB/s write speeds in a portable, high-capacity drive(1) (Based on internal testing; performance may be lower depending on host device & other factors. 1MB=1,000,000 bytes.)
- Up to 3-meter drop protection and IP65 water and dust resistance mean this tough drive can take a beating(3) (Previously rated for 2-meter drop protection and IP55 rating. Now qualified for the higher, stated specs.)
- Use the handy carabiner loop to secure it to your belt loop or backpack for extra peace of mind.
- Help keep private content private with the included password protection featuring 256‐bit AES hardware encryption.(3)
- Easily manage files and automatically free up space with the SanDisk Memory Zone app.(5). Non-Operating Temperature -20°C to 85°C
That broad definition creates an important category problem. PostgreSQL can search vectors through the pgvector extension. Search platforms can combine vector similarity with text, structured fields, ranking, and online inference. Embedded libraries can provide durable local retrieval without running a separate server. All of these can be valid choices even though they are not the same kind of product.
FAISS is the boundary case. It is a widely used approximate-nearest-neighbor and similarity-search library, but it is not a complete database with the same built-in persistence, metadata filtering, replication, or operational model as the systems compared here. It can be an excellent component inside an application, but calling it a direct substitute for a production vector database can obscure the real engineering work around it.
The comparison below therefore evaluates architecture as well as search performance.
Open-source vector databases at a glance
| Project | Deployment shape | Core data model | Most useful differentiator | Best starting point |
|---|---|---|---|---|
| pgvector | PostgreSQL extension; typically a relational server | PostgreSQL tables, rows, columns, and indexes | SQL, transactions, joins, and existing PostgreSQL operations | Applications already centered on PostgreSQL |
| Chroma | Developer-friendly local or application-integrated retrieval layer | Collections containing documents, embeddings, and metadata | Simple collection and query abstraction with metadata and document filtering | Prototypes, local applications, and small RAG systems |
| LanceDB OSS | Embedded and in-process | Tables based on the Lance data format | Vector, full-text, hybrid, SQL, and multimodal retrieval close to the application | Embedded, notebook, batch, and multimodal workloads |
| Qdrant | Dedicated single-node or distributed vector service | Collections of points with vectors and payload metadata | Payload-aware filtering integrated with vector-search traversal | Production semantic search and metadata-heavy retrieval |
| Weaviate | Dedicated vector database service | Objects stored together with vectors | Object-centric modeling, hybrid search, filtering, and reranking | AI search applications needing integrated retrieval features |
| Milvus | Lite, Standalone, or Kubernetes-oriented Distributed deployment | Collections with broad vector and scalar field types | Distributed scale, multi-vector retrieval, sparse search, and full-text capabilities | Large or heterogeneous retrieval systems |
| Vespa | Distributed search and serving platform | Documents, structured data, tensors, and ranking logic | Search, recommendation, ranking, inference, and vector retrieval in one serving system | Complex search and online-serving architectures |
This table is a map, not a performance ranking. A system that is ideal for a permission-filtered PostgreSQL application may be a poor choice for a billion-vector, multi-region retrieval service. Conversely, a distributed platform may be unnecessary complexity for a small internal RAG tool.
The practical decision tree
- Is PostgreSQL already central to the application? Start with pgvector. Keeping embeddings beside users, documents, permissions, billing records, and application transactions often eliminates an entire synchronization problem.
- Do you want retrieval inside the application rather than a network database? Evaluate Chroma and LanceDB OSS. Choose LanceDB when table-oriented, full-text, hybrid, or multimodal data handling is important.
- Do you need a dedicated vector service with metadata filtering as a first-class concern? Evaluate Qdrant.
- Do you need objects, keyword search, hybrid retrieval, filters, reranking, and RAG-oriented features in one system? Evaluate Weaviate.
- Do you need distributed operation, broad vector types, high ingest, or multi-vector retrieval at large scale? Evaluate Milvus, provided the infrastructure is justified.
- Is vector search only one component of ranking, recommendation, personalization, structured search, or online inference? Evaluate Vespa rather than treating the problem as simple embedding storage.
This is a shortlist heuristic, not a substitute for testing. The most important branch is often the first one: whether the vector index should share a transactional system with the rest of the application.
1. pgvector: the default when PostgreSQL is already central
pgvector is an open-source PostgreSQL extension, not a separate database server. It adds vector types and similarity-search capabilities while allowing embeddings to live in ordinary PostgreSQL tables alongside application data.
What it does well
- Exact search: exact nearest-neighbor search is available by default and is useful when the dataset or query volume is modest, or when predictable recall matters more than approximate-search speed.
- Approximate search: HNSW and IVFFlat indexes provide different speed, recall, memory, and build-time trade-offs.
- Relational behavior: SQL queries, joins, transactions, PostgreSQL permissions, backups, monitoring, and existing operational practices remain relevant.
- Application consistency: a document update and its embedding or metadata can be managed within the same relational architecture rather than synchronized across two databases.
HNSW generally offers a better speed-versus-recall trade-off than IVFFlat, but it requires more memory and takes longer to build. IVFFlat usually builds faster and uses less memory, but its speed-versus-recall trade-off is often weaker. The right choice depends on the index parameters, corpus, distance metric, query distribution, and hardware.
The filtering catch
PostgreSQL gives pgvector unusually powerful relational filtering, but approximate vector indexes do not make every filtered query automatically efficient. With approximate-index scans, the vector index can produce candidate rows first and the ordinary filter can be applied afterward. A restrictive predicate may therefore leave too few qualifying results even when many matching rows exist elsewhere in the table.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minutePossible responses include expanding the approximate scan, using iterative index scans where supported by the relevant pgvector version, adding ordinary or partial indexes for the filter, partitioning data when the tenant or category boundaries justify it, or using exact search for selective cases. These are schema and query-planning decisions, not just vector-index settings.
This matters for multi-tenant applications, permission-aware retrieval, product catalogs, geographic constraints, and time windows. A benchmark containing only unfiltered top-k queries can make pgvector look better or worse than it will in the actual application.
Best fit and limits
Choose pgvector first when your entities, permissions, metadata, and vectors already belong in PostgreSQL; when moderate-scale semantic search or RAG is the goal; or when transactional simplicity is more valuable than a specialized vector service.
It may become less suitable when the workload outgrows the relational system, when very high concurrent vector traffic competes with transactional queries, or when specialized distributed retrieval features are required. That is a workload-dependent limit, not a claim that pgvector has one fixed maximum scale.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #2
- Solid state performance with up to 800MB/s read speeds in a portable drive. (Based on internal testing; performance may be lower depending on host device, interface, usage conditions and other factors. 1MB=1,000,000 bytes.)
- Back up your content and memories on a storage solution that fits seamlessly into your mobile lifestyle.
- Take it with you on your adventures—up to two-meter drop protection means this durable drive can take a beating. (Based on internal testing.)
- Secure it to your belt loop or backpack for extra peace of mind thanks to the tough rubber hook.
- From Sandisk, a brand professional photographers trust to take on assignments.
2. Chroma: the easiest developer-first retrieval layer
Chroma uses a collection-oriented model for documents, embeddings, and metadata. Its query APIs support metadata filtering and document-text filtering, including logical conditions and array-containment-style filtering.
Its strongest advantage is developer velocity. A team building a local RAG prototype, educational project, notebook workflow, or small application can begin with a relatively direct retrieval abstraction instead of designing a full distributed data plane.
Best fit
- Local-first AI applications
- Proofs of concept and prototypes
- Small RAG services
- Applications where collection/query concepts are more convenient than SQL or a lower-level search API
What not to infer
Ease of use does not prove suitability for a large distributed production deployment. Chroma’s managed cloud documentation describes additional cloud capabilities, but those capabilities should not be treated as features of the self-hosted open-source core. Confirm the deployment, persistence, scaling, authentication, and backup behavior of the exact edition you plan to operate.
Chroma is Apache-2.0 licensed. That describes the open-source project; it does not mean every hosted service using Chroma is itself open source or offered under identical terms.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →3. LanceDB OSS: embedded, table-oriented, and multimodal
LanceDB OSS runs in process, with a deployment feel closer to an embedded database such as SQLite than to a resident network service. It is built around the Lance data format and stores embeddings, metadata, and multimodal data in the same table-oriented environment.
The feature set extends beyond basic dense-vector retrieval. LanceDB documents vector search, full-text search, hybrid search using secondary indexes, SQL querying, and multimodal data handling. That combination is especially useful when retrieval is part of a local data workflow rather than a standalone service accessed by many independent application nodes.
Best fit
- Embedded applications that should not require a separate database server
- Data-science notebooks and local experimentation
- Batch retrieval and file- or object-storage-oriented workflows
- Multimodal collections containing text, images, audio, or other associated data
- Applications needing vector and full-text retrieval in a table-oriented model
LanceDB OSS is Apache-2.0 licensed. LanceDB Enterprise is a separate commercial product, so claims about centralized operations, enterprise management, or managed scale must be tied to the edition being evaluated.
The architectural trade-off
Embedded is not the same as distributed. In-process operation can reduce network overhead and simplify local deployment, but it does not automatically provide the replication, failover, multi-node serving, or independent scaling characteristics of a dedicated distributed database. If several application servers need concurrent shared access, test that topology rather than assuming an embedded library will behave like a network service.
4. Qdrant: a focused vector service with strong filtering semantics
Qdrant organizes data into collections of points. A point can contain one or more vectors and optional payload metadata. Its documented capabilities include dense and sparse vectors, named vectors, HNSW indexing, payload indexes, hybrid retrieval, background segment optimization, and sharding for distributed deployments.
Why filtering is a major differentiator
Qdrant’s payload indexes are designed to work with the HNSW graph, allowing filtering to participate in semantic-search traversal rather than being treated only as an unrelated step before or after vector search. That design is particularly relevant when metadata is central to retrieval correctness: tenant boundaries, permissions, product facets, language, geography, user segments, or time ranges.
The practical question is not merely whether a database accepts a filter. It is whether the filter preserves useful recall and predictable result counts at the selectivities your application actually produces.
Best fit
- Dedicated semantic-search services
- Recommendation and similarity matching
- Catalog and faceted retrieval
- Permission- or tenant-aware vector search
- Teams that want a focused service with HTTP and gRPC interfaces and client libraries
Operational trade-offs
Qdrant’s distributed features do not remove the need to operate a distributed system. Self-hosted teams must plan for replicas, shard placement and movement, storage, backups, upgrades, and failure recovery. Qdrant’s cloud documentation describes automated replication and shard rebalancing in its managed environment, while self-hosted deployments require manual management of those concerns. Do not transfer managed-cloud automation assumptions to an installation you operate yourself.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Qdrant is Apache-2.0 licensed. As with the other projects here, evaluate the open-source core and any hosted offering separately.
5. Weaviate: object-plus-vector modeling and hybrid search
Weaviate stores data objects together with their vector embeddings. It supports vector search, keyword search, hybrid search, structured filters, reranking, and retrieval-augmented-generation workflows.
That object-centric model can be a good fit when the application thinks in terms of searchable objects rather than rows in an existing relational schema or bare points in a vector index.
Filtered vector search
Weaviate documents a pre-filtering approach in which an inverted index builds an allow-list of objects that satisfy the filter, and that allow-list is used during vector search. Its current documentation also describes a flat-search cutoff for highly restrictive filters and an ACORN filter strategy in current versions.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →These details are more informative than a feature-matrix checkbox saying that Weaviate supports metadata filtering. They also need version awareness: filtering strategies and thresholds are implementation details that can change, so verify the behavior of the release you intend to run.
Best fit and trade-offs
Consider Weaviate when you need object-centric storage, lexical-plus-semantic retrieval, flexible filtering, reranking, and an integrated AI-search workflow. Its broader feature ecosystem can shorten application development, but it also creates more conceptual and operational surface area than a small embedded library or minimal vector service.
For basic RAG, that breadth may be valuable or unnecessary. Decide based on the retrieval features you will actually use, not on the number of features listed in the product documentation.
6. Milvus: distributed scale and broad retrieval capabilities
Milvus offers three main deployment modes: Lite for local or prototyping use, Standalone for a single-machine production deployment, and Distributed for large-scale systems commonly operated on Kubernetes.
Recommended Free Tools
Milvus documentation presents guidance ranging from a few million vectors in Lite to tens of billions in distributed deployments. Those figures are vendor guidance, not independently verified capacity guarantees. Actual limits depend on vector dimensionality, data types, index configuration, filtering, update rate, query concurrency, hardware, replication, and the availability target.
Capability breadth
Milvus supports dense, sparse, and binary vectors; scalar and JSON fields; metadata filtering; range search; full-text and BM25 search; reranking; and multi-vector hybrid search. This breadth makes it suitable for systems that combine several retrieval signals or need more than one vector representation per entity.
Its distributed architecture separates access, coordination, worker, and storage responsibilities, with compute and storage that can be scaled independently in the distributed design. That can be valuable for large ingest and query workloads, but it increases the number of infrastructure decisions.
Infrastructure cost
A distributed Milvus deployment may involve metadata storage, object storage, and write-ahead-log or message-queue infrastructure. Standalone deployments can bundle or embed several dependencies, but the distributed topology is a larger data plane than most small applications need.
Rank #4
- NEARLY 2X FASTER THAN OUR PREVIOUS GENERATION(8) – move 1,000 high-res photos in under 60 seconds(6) with up to 2000MB/s transfer speeds(2).
- IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.
- POCKET-SIZED – fits easily in pockets and small bags.
- SPACE TO OWN YOUR AI CONTENT – speed and capacity to download your high-res clips and photo edits.
- 256-BIT AES ENCRYPTION(4) – helps keep private files secure with password protection.
Choose Milvus when scale, throughput, availability, or retrieval breadth justifies the infrastructure. Do not choose it merely because an application uses embeddings. For a small RAG service, Lite, pgvector, Chroma, LanceDB, or a focused single-node service may be a more rational starting point.
Milvus is Apache-2.0 licensed. Separate the license and capabilities of the open-source project from any hosted or commercial service built around it.
7. Vespa: a search, serving, and machine-learning platform
Vespa is a valid member of this comparison, but it should not be framed as a narrowly specialized vector database. It is a search and serving platform for vectors, tensors, text, structured data, ranking logic, and machine-learning inference at serving time.
That distinction is central to the buying decision. If the application needs personalized ranking, recommendation logic, structured retrieval, online inference, tensor operations, and vector search in one serving architecture, Vespa can address the whole problem. If the requirement is simply to store embeddings and return the nearest ten records, its breadth may create more learning and system-design work than necessary.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best fit
- Large-scale search and recommendation
- Ranking and personalization systems
- Applications combining lexical, structured, and vector retrieval
- Online machine-learning inference alongside serving
- Architectures where vector search is one stage in a larger ranking pipeline
Vespa’s code is Apache-2.0 licensed. As with all managed services, inspect the deployment and commercial model separately from the license of the project’s source code.
Filtering is more important than a simple feature checklist
Every modern candidate can be described as supporting some form of filtering, but that statement hides the behavior that determines whether filtered retrieval works well.
| System | How filtering fits into retrieval | Why it matters |
|---|---|---|
| pgvector | Uses PostgreSQL query planning and relational indexes. Approximate-index filtering may occur after the index scan; iterative scans, partial indexes, partitioning, or exact search may be needed. | Excellent integration with relational predicates, but restrictive filters need deliberate query and schema design. |
| Chroma | Metadata and document filtering are exposed through collection query APIs. | Convenient for application-level retrieval, though large-scale behavior should be tested rather than inferred from API simplicity. |
| LanceDB | Combines vector, full-text, hybrid, SQL, and secondary-index capabilities in an embedded table-oriented system. | Useful when structured and multimodal local retrieval matter as much as vector similarity. |
| Qdrant | Payload indexes can extend HNSW-style traversal so filters participate in semantic search. | Strong fit when metadata constraints are central to recall and result correctness. |
| Weaviate | Uses inverted-index-based pre-filtering, with documented strategies for highly restrictive filters. | Designed to avoid treating filtering as a late afterthought, but behavior is version-sensitive. |
| Milvus | Supports scalar predicates, sparse and dense retrieval, full-text/BM25 search, reranking, and multi-vector hybrid search. | Broad retrieval combinations suit heterogeneous or large systems, at the cost of more configuration. |
| Vespa | Treats vectors, tensors, text, structured data, and ranking as parts of one serving platform. | Useful when filtering and ranking are components of a complete search experience. |
Ask these questions before comparing raw queries per second:
- Does the filter apply before candidate generation, during graph traversal, after candidate generation, or through a relational query plan?
- What happens when only 0.1%, 1%, or 10% of records satisfy the predicate?
- Does the system still return the requested k results?
- How does filtered recall compare with unfiltered recall?
- Can tenants, permissions, regions, categories, and time windows be represented without creating an unmanageable number of indexes or partitions?
In permission-aware search, catalogs, recommendations, and multi-tenant systems, filtered recall and predictable latency can matter more than the fastest unfiltered ANN result.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Hybrid search and reranking: decide what the application actually needs
Dense vectors are good at semantic similarity, but lexical matching remains important for names, product codes, error messages, legal terms, and exact phrases. Hybrid retrieval combines semantic and lexical signals; reranking then applies a more expensive model or ranking function to a smaller candidate set.
- pgvector: supplies vector search inside PostgreSQL. Lexical search can be built using PostgreSQL’s broader text-search capabilities, but hybrid retrieval is an application and query-design choice rather than a single vector-extension feature.
- Chroma: exposes metadata and document-text filtering, making it approachable for document retrieval, but do not assume that document filtering is equivalent to a full hybrid ranking system.
- LanceDB OSS: explicitly combines vector, full-text, hybrid, SQL, and secondary-index capabilities in an embedded model.
- Qdrant: supports dense and sparse vectors and hybrid retrieval, which is useful when sparse and semantic representations need to be combined in a focused vector service.
- Weaviate: combines vector, keyword, and hybrid search with filters and reranking.
- Milvus: supports sparse and dense retrieval, full-text/BM25 search, reranking, and multi-vector hybrid search.
- Vespa: makes ranking, tensors, text, structured data, and online inference part of the serving architecture itself.
If exact terms and semantic meaning both affect relevance, test hybrid retrieval directly. A dense-only benchmark cannot establish which system will produce the best search experience.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Deployment and operations: the cost beyond the API
Embedded and in-process systems
Chroma in a local or application-integrated setup and LanceDB OSS reduce the operational burden of running a separate network service. They can be attractive for desktop tools, notebooks, local agents, batch jobs, and early prototypes.
The trade-off is shared access and operational maturity. You must validate concurrent writers, multi-process access, file or object-storage behavior, backup procedures, restart recovery, and how the application will serve the data when it runs on multiple hosts.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
- Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Relational integration
pgvector inherits PostgreSQL’s mature transactional, security, backup, and query ecosystem. That is a major advantage when the vector records and application records have the same lifecycle. It also means vector queries consume resources in the same database that may be serving critical OLTP traffic.
Separate read capacity, connection management, index-build scheduling, and resource isolation should be part of the proof of concept. A vector feature that is operationally simple can still create contention with ordinary application queries.
Dedicated vector services
Qdrant and Weaviate provide a focused service boundary. Milvus and Vespa go further toward distributed search and serving platforms. These architectures can isolate vector workloads and scale independently, but they introduce another system to secure, monitor, back up, upgrade, and integrate with application identity and data pipelines.
For distributed deployments, evaluate shard placement, replication, rebalancing, object storage, metadata services, message or write-ahead-log infrastructure, rolling upgrades, node loss, and recovery time. Kubernetes can be an appropriate foundation for large deployments, especially for Milvus Distributed, but Kubernetes itself is not a substitute for database-operational expertise.
Recommended Free Tools
Licensing: open-source core versus hosted product
Chroma, LanceDB OSS, Qdrant, Milvus, and Vespa document Apache-2.0 licensing for their open-source projects or cores. pgvector is an open-source PostgreSQL extension with its own repository license. Read the actual license and any additional notices before embedding a component in a commercial product.
Do not collapse three different questions into one:
- Can I inspect, run, modify, and redistribute the source? This is primarily a project-license and distribution question.
- What features are present in the open-source edition? Authentication, backups, multi-region behavior, management interfaces, and enterprise controls may differ by edition.
- What am I buying from a managed provider? Hosting, support, automated scaling, upgrades, recovery, and service-level commitments are commercial-service questions.
An open-source core does not make every hosted offering open source. Availability, pricing, feature packaging, and partner programs can also change, so verify them directly for the specific service and date of purchase.
How to benchmark them without producing a misleading ranking
There is no responsible universal speed ranking for these systems. Results depend on the engine, dataset, distance metric, scenario, deployment mode, server and client topology, client implementation, index settings, and hardware. Benchmark projects routinely include cases where systems fail, time out, or are not directly comparable across scenarios.
A useful benchmark should reproduce your workload, not a vendor’s preferred workload.
Measure at least these dimensions
- Recall at a stated k: compare approximate results with an exact ground-truth search using the same distance metric.
- Latency distribution: record p50, p95, and p99, not just an average.
- Throughput: state the query-per-second result and the concurrency level that produced it.
- Ingestion: measure initial load speed, sustained ingest, index-build time, and the effect of concurrent queries.
- Mutation behavior: test updates and deletes, including whether index maintenance creates latency or storage overhead.
- Filtered recall: run several predicate selectivities, such as broad, moderate, and highly restrictive filters.
- Resource footprint: record memory, disk or object-storage use, index size, and CPU consumption.
- Recovery: measure restart time, restore time, replica recovery, and behavior after node or storage failure.
- Operational labor: count the services, configuration, alerts, upgrades, and runbooks required by the chosen topology.
- Total cost: include compute, storage, network, backups, managed-service premiums, and engineering time.
Keep the experiment fair
Use the same embedding model, corpus, query set, vector dimensionality, distance metric, hardware, client concurrency, warm-up process, and target recall. Record index parameters and deployment configuration. Test both warm-cache and cold-start or restart behavior when those conditions matter to the application.
Vendor-sponsored benchmarks can help identify useful test cases, but they are not neutral cross-product proof. Treat them as input to your test plan, not as a universal leaderboard.
A practical proof-of-concept plan
- Define the retrieval contract. Write down the required k, distance metric, acceptable p95 and p99 latency, target recall, update rate, tenant model, filter predicates, and availability target.
- Choose a representative corpus. Include short and long documents, duplicate or near-duplicate content, multilingual data if relevant, and the metadata distributions that production filters will use.
- Create real queries. Use representative user questions, exact-name lookups, product or error-code searches, permission-constrained requests, and queries with no valid matches.
- Test the smallest plausible architecture first. If PostgreSQL already exists, test pgvector. If the application is local or embedded, test Chroma or LanceDB OSS. Do not begin with a distributed cluster unless the requirements demand it.
- Add filtered and hybrid cases. Measure filtered recall and result counts, not just unfiltered nearest neighbors. Test lexical-plus-vector queries if users search for proper nouns, codes, or exact phrases.
- Exercise failure and maintenance. Restart nodes, rebuild indexes, apply updates and deletes, take and restore backups, and observe behavior during ingestion and upgrades.
- Make the decision from bottlenecks. Move to Qdrant, Weaviate, Milvus, or Vespa when a measured requirement needs their filtering, retrieval breadth, distributed scale, or serving capabilities—not because the product category is fashionable.
Recommendations by workload
| Workload | First systems to evaluate | Reason |
|---|---|---|
| Existing SaaS application with users, permissions, and transactions in PostgreSQL | pgvector | One relational system can hold application data, metadata, permissions, and embeddings. |
| Local RAG prototype or educational project | Chroma; also LanceDB OSS | Low-friction collection or embedded workflows let the team validate retrieval quickly. |
| Embedded or notebook retrieval over multimodal data | LanceDB OSS | In-process operation and vector, full-text, hybrid, SQL, and multimodal table capabilities fit the workflow. |
| Production semantic search with strict tenant, catalog, or permission filters | Qdrant; compare with pgvector if PostgreSQL is already central | Filtering behavior and operational topology deserve direct testing. |
| Object-centric AI search with keyword, vector, hybrid, and reranking features | Weaviate | Its data model and retrieval features align with integrated object search. |
| Large-scale or multi-vector retrieval with broad scalar and text capabilities | Milvus | Distributed modes and broad retrieval support can justify a larger infrastructure footprint. |
| Recommendation, personalization, ranking, and online inference | Vespa | Vector retrieval can be designed as one part of a complete serving and ranking platform. |
Common mistakes to avoid
- Choosing by unfiltered ANN speed alone. Your production workload may be dominated by permissions, tenants, categories, or time ranges.
- Calling a managed service open source. Identify the source project, the self-hosted edition, and the hosted service separately.
- Assuming embedded means production-ready for every topology. Validate shared access, recovery, concurrency, and multi-host serving.
- Starting with a distributed cluster for a small application. Milvus Distributed, Vespa, and other cluster-oriented architectures can be excellent at their intended scale and excessive at prototype scale.
- Ignoring the existing database. A dedicated vector service may improve isolation, but it also introduces data synchronization, identity, backup, and operational work.
- Assuming all filters are equivalent. Pre-filtering, integrated graph traversal, post-filtering, and relational planning can produce very different recall and latency results.
- Confusing document filtering with hybrid ranking. Filtering documents by text is not necessarily the same as combining lexical and semantic relevance scores.
- Benchmarking with synthetic data only. Real metadata distributions and real queries often expose the limitations that a clean unfiltered benchmark misses.
Frequently Asked Questions
Is FAISS a vector database?
FAISS is primarily an approximate-nearest-neighbor and similarity-search library. It can be an excellent search component, but it does not provide the same complete database model, metadata filtering, persistence, replication, and operational tooling as the vector databases compared here.
Should I use pgvector or a dedicated vector database?
Start with pgvector when PostgreSQL already owns the application data, permissions, and transactions. Consider Qdrant, Weaviate, Milvus, or Vespa when measured requirements call for an independent vector service, specialized filtering, broader hybrid retrieval, distributed scale, or a full search-and-serving platform.
Which open-source vector database is best for RAG?
There is no single best choice. Chroma is a practical developer-first starting point for local or small RAG projects; pgvector fits RAG applications already using PostgreSQL; LanceDB OSS suits embedded or multimodal RAG; and Qdrant or Weaviate may be better when production filtering and integrated retrieval features are central.
Does an Apache-2.0 vector database make its hosted service open source?
No. The Apache-2.0 license applies to the relevant open-source project or core. A hosted service may add proprietary software, automation, support, commercial features, and different terms. Evaluate the self-hosted source and managed offering separately.
The Bottom Line
Bottom line: pick the architecture before you pick the product. Use pgvector for relational simplicity, Chroma for a straightforward local start, LanceDB OSS for embedded multimodal retrieval, Qdrant for focused filtered vector search, Weaviate for object-centric hybrid retrieval, Milvus for distributed scale, and Vespa for full search and serving. Then validate the choice with a proof of concept that measures filtered recall, tail latency, recovery, resource use, and operational effort—not just unfiltered top-k speed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

