iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Re-embed your documents with the new model before using its query embeddings for retrieval. Stored vectors are representations produced by a particular embedding model and configuration; equal dimensions do not make vectors from different models interchangeable. Keep the old retrieval path available while you rebuild and validate the new one, then switch production queries to the new representation.
Why the old vectors stop being a safe match
Vector search compares positions in an embedding space. A model change can alter how text is represented and how similarity relates to relevance. If queries are embedded with the replacement model while documents retain vectors from the previous one, nearest-neighbor results may no longer be meaningful.
Matching dimensions do not establish compatibility: two models can output vectors with the same number of elements without preserving the same semantic geometry or relevance behavior. MongoDB’s Voyage AI migration documentation says to regenerate the entire corpus so stored and query vectors come from the same model, even when dimensions and element type match. MongoDB’s migration guidance
Recommended Free Tools
What to do before switching models
Recover the text that produced the vectors
Confirm that you can retrieve the original documents or the exact chunks that were embedded. A vector is not a substitute for its source text, so vectors alone cannot be used to generate the replacement representation. If source content is missing, identify what can be reconstructed before planning the backfill.
#1 Best Overall
Pin the model and record its configuration
Choose the successor model and record its identifier, version where available, dimensions, and relevant embedding configuration alongside the new representation. OpenAI’s backward-compatibility guidance recommends pinned model versions for more consistent behavior; it addresses API and model behavior, not compatibility between embeddings from distinct models. OpenAI’s backward-compatibility guidance
Check the target schema and index
Verify that the database’s vector field or index supports the successor model’s output dimensions and that your deployed database version supports the migration approach you plan to use. For MongoDB’s documented automated-embedding path, the new index dimension must match the successor model’s output. Database capabilities and migration steps vary by product and deployment.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
A migration sequence that limits cutover risk
- Choose a migration path. Decide whether to build a new collection or index, add another vector representation alongside the old one where supported, or use a managed embedding workflow. Keep the existing retrieval path available if your system permits it.
- Backfill from source text. Generate new embeddings for the retained documents or chunks and write them to the new field, vector, or index. Batch and monitor the work according to your provider’s and database’s limits. MongoDB notes that regeneration through its automated workflow incurs additional embedding costs, but does not specify a universal amount.
- Keep incoming changes in sync. During the backfill, ensure new and changed documents also receive the new representation. Use the database’s documented migration mechanism so that the new path does not fall behind the corpus.
- Evaluate on your application’s data. Run a representative query set against the new representation. Inspect whether relevant results are retrieved and ranked appropriately for your use case. The cited guidance does not establish a universal acceptance threshold, cost, latency, or quality improvement; set criteria that reflect your application.
- Switch reads after validation. Route production queries to the new representation only when it is populated and has passed your checks. Keep the old path available for rollback until the new one is stable.
- Retire the old representation deliberately. Remove the old field, vector, or index only after production validation and when you no longer need the rollback option.
Migration options documented by database vendors
| Approach | Documented option | Trade-offs |
|---|---|---|
| New collection or blue-green index | Qdrant documents migrating to a separate collection when the named-vector method is unavailable. Qdrant migration guide | Provides a separate representation and rollback boundary, but requires parallel data and index work. Infrastructure cost depends on the deployment. |
| Additional named vector | For collections created with named vectors on Qdrant version 1.18 or later, the guide describes adding the successor as another named vector, backfilling in the background, switching the query’s using parameter, and then deleting the old vector. Qdrant migration guide |
Can avoid a second collection in eligible cases, but depends on the collection schema and deployed version. |
| Managed automated embedding | MongoDB’s Voyage AI guide describes changing the model in the managed index configuration and allowing the service to regenerate the index and embeddings. It says queries can continue against the old index definition while rebuilding. MongoDB migration guide | Reduces application-managed embedding work, but requires checking schema and deployment constraints; regeneration incurs additional embedding costs. |
| Self-managed embedding | MongoDB’s guide also describes regenerating from corpus data through a self-managed path and writing the new vectors to a suitable field and index. MongoDB migration guide | Offers control over backfill and validation while leaving your application responsible for consistency, retries, and cutover. |
| New vector field | Zilliz Cloud’s runbook describes adding a new vector field, migrating existing and incoming records, validating the new representation, and moving production search to that field. Zilliz migration guide | Follows that service’s documented workflow; confirm the steps and capabilities for your own product and deployment. |
Choose the approach that fits your system
No single migration method is best for every deployment. Compare the options against the constraints that affect your cutover:
Quick Recap
Best Value
Rank #4
Rank #3
- Continuity and rollback: Can the old index keep serving queries while the new representation is built? Can you route reads back if validation fails?
- Corpus and backfill: How much source text must be embedded, and how long can the rebuild take within your provider and database limits?
- Schema and version: Does the current collection or index support a second vector, or must you create a separate one?
- Source availability: Can you reconstruct the original text and keep new or updated records synchronized?
- Cost and evaluation: What will the provider charge for your actual backfill, and can you evaluate relevance before routing production traffic?
Common mistakes to avoid
- Switching only the query encoder: That leaves queries and stored documents in representations produced by different models.
- Treating equal dimensions as proof of compatibility: Dimensions describe vector size, not whether different models produce comparable representations.
- Re-embedding only a sample for production: A sample can support evaluation, but the production corpus needs its new representation for consistent retrieval.
- Deleting the old path too early: Removing it before the new path is validated takes away a practical rollback option.
- Assuming a vendor workflow applies everywhere: The documented mechanics depend on the database, schema, deployment mode, and—in Qdrant’s named-vector method—version.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

