Vector Databases for AI: How They Work and When to Use Them
How vector databases work and how to choose one: approximate nearest neighbour indexes such as HNSW, filtering, hybrid search, scaling, cost, and when Postgres with pgvector is enough.
Quick answer
A vector database stores embeddings with metadata and quickly finds the vectors nearest to a query, using approximate nearest neighbour indexes such as HNSW or IVF. It powers semantic search and RAG retrieval. If your data already lives in Postgres and scale is moderate, pgvector is often enough; dedicated vector databases suit large scale or advanced vector features; search engines suit teams that need strong keyword and hybrid search. Choose by filtering needs, scale, hybrid search, operations and cost, tested on your own data.
Where This Fits
Vectors come from embedding models. Retrieval quality usually improves with hybrid search and reranking. The full pipeline is in the RAG guide, and product search use is in ecommerce semantic search.
How Vector Search Works
Each item (a document chunk, product or image) is stored as a vector: a list of numbers produced by an embedding model. A query is embedded the same way, and the database returns the stored vectors with the smallest distance (cosine, dot product or Euclidean). Comparing against every vector is exact but slow at scale, so databases build approximate nearest neighbour (ANN) indexes that find near-best matches quickly.
| Index type | How it works | Trade-offs |
|---|---|---|
| Flat (exact) | Compare with every vector | Exact, slow at scale |
| HNSW | Multi-layer proximity graph | Fast and accurate; more memory, slower builds |
| IVF | Cluster vectors, search nearest clusters | Less memory; needs training and tuning |
| Quantization | Compress vectors | Lower memory and cost; some accuracy loss |
Filtering and Multi-Tenancy
Real queries filter: this customer's documents, this product line, content the user may see. With approximate indexes, filters applied after the index scan can return too few results. Systems handle this differently: pre-filtering, filtered index traversal or scanning further. pgvector added iterative index scans in version 0.8.0 for this reason. For multi-tenant products, decide between shared indexes with tenant filters and separate indexes per tenant based on isolation needs and scale.
Choosing the Type of System
Postgres and pgvector
pgvector adds vector types and indexes to Postgres. You get transactions, joins with business tables, SQL filtering and your existing backups and operations. It supports HNSW and IVFFlat indexes, half-precision vectors (halfvec) for smaller indexes and iterative scans for filtered queries. It suits many RAG and semantic search systems; at very large scale or with heavy query loads, tuning and dedicated infrastructure become more important.
Choosing a vector store for your AI application?
ZSpace Labs can benchmark pgvector, dedicated vector databases and search engines on your own data and queries before you commit.
Dedicated Vector Databases and Search Engines
Dedicated vector databases such as Qdrant, Pinecone, Weaviate and Milvus focus on vector workloads: scaling, filtering, quantization and often hybrid search with sparse vectors. Search engines such as Elasticsearch and OpenSearch combine mature keyword search with vector search and rank fusion, which suits content-heavy retrieval. Each adds a system to operate or a managed service to pay for.
Selection Criteria
- Where your source data already lives
- Number of vectors now and in two years, and dimensions
- Filtering complexity and permission requirements
- Need for keyword or hybrid search
- Latency and throughput targets
- Managed service versus self-hosting, and data residency
- Team familiarity and operational tooling
- Total cost including replicas and re-indexing
Operations and Cost
Plan for backups, re-indexing when embedding models change, index rebuild times, memory usage (HNSW indexes can be large), replicas for availability and monitoring of recall and latency. Reduce cost with smaller embedding dimensions where quality allows, quantization, and removing stale vectors.
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Fast semantic search at scale | Approximate results; recall must be measured |
| Find similar items without exact terms | Weak on exact codes without hybrid search |
| Metadata filtering and multi-tenancy | Filtering and ANN interact in tricky ways |
| Managed options reduce operations | Another system and cost to manage |
How to Choose Step by Step
- 1. Write down data size, filters and latency needs
- 2. Shortlist pgvector, one dedicated database and one search engine
- 3. Load a realistic sample with your embeddings and metadata
- 4. Run your real queries and measure recall, latency and filtered results
- 5. Estimate cost at expected scale
- 6. Decide, documenting re-indexing and backup plans
Example: Vector Search in Postgres With pgvector
For teams already on Postgres, a minimal setup looks like this. Check the pgvector documentation for current syntax and tuning parameters.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE doc_chunks (
id bigserial PRIMARY KEY,
tenant_id uuid NOT NULL,
source_id text NOT NULL,
content text NOT NULL,
embedding vector(1024) NOT NULL
);
CREATE INDEX ON doc_chunks USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON doc_chunks (tenant_id);
-- with filtered queries, consider iterative index scans (pgvector 0.8.0+)
SET hnsw.iterative_scan = relaxed_order;
SELECT id, source_id, content
FROM doc_chunks
WHERE tenant_id = $1
ORDER BY embedding <=> $2 -- cosine distance to the query embedding
LIMIT 20;Sizing and Capacity Planning
Estimate vectors (chunks per document times documents, including growth), dimensions and precision to size storage and memory. Graph indexes such as HNSW perform best when they fit in memory. Reduce footprint with shorter embeddings where quality allows, half-precision or quantized vectors, and removing stale content. Plan capacity for re-indexing, which can temporarily double storage, and for query peaks. Load-test with realistic filters, not just unfiltered nearest-neighbour queries.
Worked Example
An illustrative scenario, not a client case: a SaaS company adds document Q&A for customers. Its data is already in Postgres, with a few million chunks across tenants. pgvector with an HNSW index, tenant filters and iterative scans meets latency targets in testing, avoiding a new system. The team documents a threshold at which it would revisit a dedicated vector database.
Common Mistakes
- Choosing from benchmarks that do not match your filters
- Ignoring filtered-query recall
- No plan for re-embedding
- Vector-only retrieval for content full of identifiers
- Over-provisioning dimensions and replicas
Need retrieval infrastructure that scales sensibly?
Talk to ZSpace Labs about RAG and vector search development and database and backend architecture.
Conclusion
Vector databases make semantic retrieval fast, but the right choice depends on your data, filters and team. Start with what you already run, measure filtered recall and add specialised systems only when needed. Related: vector embeddings, hybrid search and RAG.
Common questions
A database designed to store vector embeddings and find the vectors most similar to a query vector quickly, usually with approximate nearest neighbour indexes, along with metadata filtering.