Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To shrink vectors in a pgvector-backed Spring AI application, you shorten the embedding output and set the PgVectorStore dimensions to the same value. The embedding model’s output width, the vector column width, and the index all have to agree. Pick the width first, because it sets the storage cost, the index limit you are working under, and the amount of re-embedding you will need to do. Then confirm retrieval quality on your own queries before you switch production traffic.

Why vector width matters for the index

Each stored embedding is a fixed-length array of floats, and its length is its dimension count. A 1536-dimension vector holds 1536 numbers, so reducing the width reduces the bytes each row needs and the size of the index built over those rows. The saving is real but proportional: a table with 10 million 1536-dimension rows and a 512-dimension version of the same rows will differ in size, but this article does not claim a fixed percentage for any workload, because none was measured here.

The more immediate reason to slice is the index ceiling. In pgvector, the vector type can be indexed only up to a fixed number of dimensions. If your embedding model produces wider vectors than that limit, you cannot build an HNSW index on the plain vector column without changing either the width or the type.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The limits you are working against

Two documented limits matter for this decision. The Spring AI PgVectorStore reference cites 2000 dimensions as the maximum for pgvector HNSW indexes in its vector example. The pgvector project README documents the vector type up to 2000 dimensions and the halfvec type up to 4000 dimensions.

Option Documented dimension ceiling What changes in your setup What is not established here
Keep the model’s default width on vector 2000 (pgvector README; Spring AI PgVectorStore reference for HNSW) Nothing in the model or column; works only if the default width is at or below 2000 Whether your chosen model’s default width exceeds the ceiling; check the model’s documentation
Request a shorter output from a model that supports it (for example OpenAI text-embedding-3 with the dimensions parameter) Set by the width you request, which must still fit the index type you use Re-embed all documents, set the same width in PgVectorStore, and recreate the table Retrieval quality change for your corpus; not established by these sources
Use the halfvec type where your integration supports it 4000 (pgvector README) Column type changes; Spring AI’s reference does not show it switching the column type automatically, so this needs its own schema work Whether your Spring AI version and index setup support halfvec for this store; not stated in the Spring AI reference

The Spring AI reference uses vector(1536) as its example column. That is an example width, not a universal model width. Do not assume the store will switch the column type to halfvec because your width is above 2000.

Step 1: Choose a target width

Work backward from three constraints:

  • Index type. If you need an HNSW index on the vector type, the target must be at or under 2000 dimensions.
  • Model support. The width must be one the embedding model can actually produce. For OpenAI text-embedding-3 models, the dimensions request option is documented as supported in text-embedding-3 and later models.
  • Quality floor. Shorter is not automatically acceptable. Pick two or three candidate widths you can measure, not one guess.

Keep the chosen width fixed for the life of the table. A vector stored at one width cannot be compared with a query vector at another width.

Step 2: Shorten the embeddings at the model

For OpenAI embeddings, the shortening happens in the embedding request, not in the database. The OpenAI announcement of its text-embedding-3 models, dated January 25, 2024, introduced the shortening approach, and the embeddings API reference documents the dimensions parameter for text-embedding-3 and later models. Source links are at the end of this article.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In Spring AI, set the OpenAI embedding model’s dimensions option to your target width. The exact property or option name depends on the Spring AI version your project uses. Check the OpenAI embedding section of the Spring AI reference for the version you pin, and confirm the name there rather than copying it from an older example.

OpenAI states that its API embeddings are L2-normalized by default, including after shortening, and that cosine similarity and Euclidean distance produce identical rankings for normalized vectors. That statement applies to OpenAI embeddings specifically. If you switch to a different embedding provider, check its normalization behaviour before assuming the same ranking equivalence.

Step 3: Set the PgVectorStore dimensions to match

The PgVectorStore dimensions setting controls the width of the embedding column when the table is created. Set it to the same value as the model’s output width:

spring.ai.vectorstore.pgvector.dimensions=1024
spring.ai.vectorstore.pgvector.initialize-schema=true

The value 1024 is only an illustration. Use your chosen width.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • If you omit dimensions, PgVectorStore reads the width from the configured EmbeddingModel. That is convenient, but it means a model change can silently change the expected width, so an explicit value is safer in production.
  • initialize-schema defaults to false. The reference says schema initialization must be explicitly enabled. Setting it to true creates the table if it does not exist; it does not reshape a table that already exists.

Step 4: Migrate an existing table

Changing dimensions does not change an existing table. The Spring AI reference states that the setting affects the column at table creation and that changing it requires recreating the vector_store table. Plan the migration as a rebuild:

  1. Confirm the target width and the new embedding configuration in a staging environment.
  2. Stop or pause writes that use the old embedding model, so no vectors of the old width are added during the rebuild.
  3. Drop the existing vector_store table (or create a new table name so the old one stays available for rollback).
  4. Set dimensions to the new width and let initialize-schema create the table, or create it yourself with a vector(N) column that matches N.
  5. Re-embed every document with the new model configuration and load it into the new table. Old vectors cannot be reused at a different width.
  6. Run your retrieval evaluation against the new table before routing queries to it.

If the widths differ, the failure appears at insert or query time as a dimension mismatch error from pgvector, not as a silent quality drop. That makes a mismatch easy to catch in staging, but only if you test the full path: embedding, insert, and similarity search.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Validate retrieval on your own corpus before switching

Shortening trades size for fidelity. The OpenAI announcement reports that text-embedding-3-large shortened to 256 dimensions outperformed unshortened text-embedding-ada-002 at 1536 dimensions on the MTEB benchmark. That is one benchmark comparison between two specific models and widths. It does not show that your application will keep the same answer quality at a given width.

Before you switch production traffic, build a small evaluation set from real questions your users ask, with the documents each answer should retrieve. Then compare the current and candidate configurations on:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Recall or task-level answer quality for those queries, at the same top-k you use in production.
  • Stored vector size and index size, measured on your table, not estimated from the width alone.
  • Search latency under your expected load.
  • Build and update cost for the index, including the re-embedding time and API cost for your corpus.

Keep the old table until the new one passes these checks. If the candidate width loses too much recall on your queries, the cheaper fallback is usually a wider model width within the index limit, not a different index type chosen without testing.

Where to run it

If you run PostgreSQL yourself, install pgvector and enable it with CREATE EXTENSION vector; in the target database. If you use a managed PostgreSQL service, confirm that the service offers pgvector and the version you need before you design the schema, since extension availability varies by provider and version. No specific provider is recommended here.

What this does and does not solve

Slicing embeddings reduces the width of each stored vector and can keep an HNSW index within the documented vector ceiling. It does not guarantee unchanged retrieval quality, and it does not replace index tuning, query design, or chunking decisions. Treat width as one lever, and measure it against the other levers on your own workload.

Official references used in this article: the Spring AI PgVectorStore reference, the pgvector project README, the OpenAI announcement of new embedding models and API updates, the OpenAI embeddings API reference, and the OpenAI embeddings FAQ.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.