
What Is a Vector Database? A Practical Guide for RAG Builders
TLDR: A vector database stores embeddings, which are arrays of numbers that represent the meaning of text, and finds the nearest ones to a query embedding fast enough to use at inference time. You need one when your RAG corpus is private documents and your queries are semantic. You do not need one when the corpus is the live web, because a search API retrieves the sources directly. This guide covers how vector search works, when the database earns its place, and the production failures it introduces.
Retrieval-augmented generation needs a retriever, and the vector database is the most commonly assumed one. The assumption is worth questioning, because a vector database is an indexing commitment you build and operate, not a library you import. If your questions are answered by the current web, You.com provides retrieval directly through the Web Search API, and the vector database question never arises. This guide is for the other case: a corpus you own, such as product documentation, contracts, or support tickets, where you decide what gets indexed and how.
What Is a Vector Database?
A vector database is a storage and index system for embeddings, with a query interface that returns the closest vectors to a given input vector. The embedding part is the model that turns a piece of text into a fixed-length array of numbers, and the database part is the structure that makes searching those arrays tractable at scale.
The reason a dedicated system exists at all is that nearest-neighbor search over raw arrays is quadratic in the worst case and linear per query at best, which is too slow for inference-time retrieval over a large corpus. Vector databases answer this with approximate nearest neighbor indexes, which trade a small amount of recall for orders of magnitude faster queries. The most widely used index structure is Hierarchical Navigable Small World graphs, introduced by Malkov and Yashunin in the paper Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs.
How Does Vector Search Work?
At query time the flow is fixed: embed the query with the same model that embedded the corpus, search the index for the nearest stored vectors, and return the text chunks those vectors represent. Closeness is measured with a distance function such as cosine similarity, and the query embedding is computed against the query in isolation.
Two properties follow from this design and explain most of the failure modes later in this guide. First, similarity is between embeddings, not between meanings, so the embedding model is a single point of failure for the whole system. Second, the index is approximate, so a query can miss a nearby vector, and the miss rate is a tunable tradeoff against speed rather than a bug.
When Do You Actually Need One?
The decision has three branches, and the honest answer for many teams is the third one.
- Private corpus, semantic queries: yes. Contracts, internal policies, and product documentation where users ask questions in their own words are the canonical case, because paraphrase is the norm and keyword search misses it.
- Live web corpus: no. The web is already indexed, and a search API with a content extraction step retrieves fresher sources than any index you can maintain. Our RAG with web search guide covers that architecture.
- Small corpus, exact-term queries: probably not yet. A corpus that fits in memory and queries that use the same vocabulary as the documents, such as error codes and part numbers, are served by keyword search, which is exact where vector search is approximate. The tradeoffs are laid out in our vector search vs keyword search comparison.
When you do need one, the retrieval layer is one stage of a larger pipeline. Our RAG pipeline guide covers the surrounding stages, including chunking, which determines what the vectors represent. The chunking decision matters as much as the database choice, and our chunking strategies guide compares the options.
What Are the Deployment Options?
Vector storage spans a range from an extension on a database you already run to a fully managed service, and the choice is mostly an operational tradeoff.
In your existing database. pgvector adds vector columns and approximate index queries to PostgreSQL, so the vector store lives beside your relational data with one system to operate. This is the pragmatic default for teams already running Postgres, at corpus sizes into the millions of vectors.
Dedicated engines. Purpose-built vector databases offer richer index tuning and higher throughput at scale. They also add a second stateful system to operate, with its own backups, scaling, and failure modes. The tradeoff pays when query volume or corpus size outpaces the extension path.
Managed services. Hosted vector search removes operations and adds a vendor dependency and data residency considerations. For teams without database operations capacity, this is often the fastest path to a working system, at the cost of control.
The choice between them is an infrastructure decision, and the retrieval quality above them is determined by the embedding model and the corpus preparation, not the engine.
What Fails in Production?
The embedding model swap. Change the embedding model and every stored vector is from a different space than the query vectors. Similarity scores become meaningless, and retrieval quietly degrades to near-random while nothing raises an error. Detection: store the embedding model name with every vector, refuse queries whose model name does not match, and re-embed the corpus on any model change rather than mixing.
The recall cliff. Approximate indexes have a speed-recall knob, and settings tuned for one corpus size or access pattern can starve another. A filtered query, such as tenant equals a customer, that matches few rows can defeat the graph index and return too few candidates even though matches exist. Detection: run a labeled query set on a schedule and track recall against exact search on a sample. A recall drop after a settings change or a corpus growth spurt is an index problem, not a model problem.
The drift between index and source. The corpus changes, re-embedding jobs partially fail, and the index holds stale chunks mixed with fresh ones. Answers cite removed policies. Detection: version the index with a corpus fingerprint, count chunks per source document at re-embed time, and alert when a document's chunk count changes by more than the edit could explain.
The paraphrase miss. Vector search is the fix for paraphrase, but it is not a guarantee. Slang, abbreviations, and domain jargon that never appear in the corpus embed away from the queries that use them. Detection: log the queries that retrieved nothing above the similarity threshold, and review a sample. A pattern of empty retrievals from one user population is a vocabulary gap, and the fix is usually query expansion or a hybrid with keyword search rather than a new model.
Related Guides
- Vector Search vs Keyword Search for RAG Pipelines
- How to Build a RAG Pipeline: A Step-by-Step Guide for Developers
- 5 RAG Chunking Strategies in 2026: Fidelity, Cost, and Complexity
- What Is an API for RAG? Retrieval-Augmented Generation for Modern AI Applications
- Web Search API guide
FAQ
Do I need a vector database for web RAG? No. Retrieval over the live web is what a search API does, with the search engine as the index and result pages as the corpus. A vector store enters the picture only if you re-index extracted web content for repeated semantic queries.
Can I use PostgreSQL as my vector database? Yes. The pgvector extension adds vector columns and approximate nearest neighbor indexes to Postgres, and it is a common production choice for corpora into the millions of vectors, with the operational advantage of one database to run.
What happens if I change embedding models later? The stored vectors and the new query vectors live in incompatible numeric spaces, so retrieval quality collapses without any error. Store the model name with the vectors, block mismatched queries, and re-embed the full corpus when you switch.
Why did my filtered searches start returning fewer results? Highly selective filters can shrink the candidate set an approximate index sees, which starves the search even when matches exist. This is a known behavior of filtered approximate search rather than a data loss bug, and the fixes are pre-filtering strategies or hybrid retrieval.
LI Test
LI Test
Share Article:
Related resources.

6 Local AI Models You Can Run Today: Sizes, Context, and Licensing
September 4, 2026
Blog

Context Window: Meaning and Optimization Tips
May 26, 2026
Blog

What Is Semi Structured Data: A Developer's Guide
May 4, 2026
Blog

What Is a Web Crawler in a Website and How Does It Differ From a Search API?
February 11, 2026
Blog

.png)