When I started building the RAG pipeline for my support-automation system, I spent too much time reading benchmark comparisons. Every article had tables of impressive numbers — P95 latencies, QPS figures, memory usage per million vectors. What I discovered after shipping to production is that those benchmarks almost never predict real-world behavior in your specific setup.
What actually matters is: how does this database fit into your infrastructure? How does it fail? What happens when you need to update embeddings at 2am? Those questions don't show up in comparison tables.
This post shares what I learned from building and operating a production vector search system — the RAG backbone of a support-automation pipeline handling hundreds of tickets per day.
The System I Built Against
The support-automation system I'll reference throughout is a RAG pipeline that:
- Ingests and re-embeds support documentation on a nightly schedule
- Serves semantic search results to a triage agent that routes and drafts ticket responses
- Handles spiky load patterns (tickets cluster around business hours)
That context matters because my choices were shaped by it. A system doing real-time image similarity search would land on different answers.
Understanding What Vector Databases Actually Do
Before comparing options, it helps to be precise about what a vector database provides that a regular database doesn't.
When you embed a piece of text with a model like claude-haiku-4-5-20251001 or a dedicated embedding model, you get a high-dimensional float array — typically 256 to 3072 dimensions depending on the model. A vector database stores these arrays and answers the question: "given this query vector, which stored vectors are most similar?"
The core operation is approximate nearest neighbor (ANN) search. "Approximate" is the key word — for large collections, exact nearest-neighbor search is too slow, so these systems trade a small amount of recall accuracy for large speed gains using indexing structures like HNSW (Hierarchical Navigable Small World graphs).
Here's a minimal embedding + search pattern using pgvector, which is what my support-automation system runs on:
The <=> operator is pgvector's cosine distance operator. The index does the heavy lifting to avoid a full table scan on every query.
The Options I Evaluated
When I was building the support-automation RAG pipeline, I seriously evaluated four options. Here's my honest assessment of each — based on operational experience rather than synthetic benchmarks.
pgvector
pgvector is a PostgreSQL extension that adds a vector type and ANN index support. I chose it for the support-automation system and have no regrets.
Why it worked for me: The support-automation system already ran Postgres. Adding pgvector meant one fewer service to operate, one fewer failure mode, and the ability to do hybrid queries joining ticket metadata with semantic search in a single SQL statement.
Where it struggles: At very large scale (tens of millions of vectors), pgvector's HNSW index has memory constraints. The index lives in shared memory, and you need to tune maintenance_work_mem carefully during index builds. For my scale (hundreds of thousands of docs), this was never a real issue.
When to choose it: You already run Postgres, your collection fits on a single node, and you want transactional consistency between your vector data and your relational data.
Pinecone
Pinecone is a fully managed, purpose-built vector database. It abstracts away all infrastructure.
What it gets right: The operational simplicity is real. There's no cluster to manage, no index tuning knobs to turn, no memory sizing to get wrong. If you're a small team that wants to ship quickly and doesn't want to think about vector database operations, Pinecone delivers on that.
Where it struggles: The managed service model means you're dependent on Pinecone's availability and pricing. Metadata filtering has limitations — you can't do arbitrary SQL-style joins. And if you need to do anything outside the core search API, you're working against the grain.
When to choose it: You want operational simplicity above all, you don't need to join vector search with relational queries, and the pricing model works for your scale.
Qdrant
Qdrant is a purpose-built vector database written in Rust, available as both self-hosted and managed cloud.
I evaluated Qdrant seriously because its filtering capabilities are more expressive than Pinecone's. It supports complex filter trees with must/should/must_not conditions, which maps well to the kind of structured metadata filtering my triage agent needs.
Where it struggles: It's another service to operate. If you're not already containerizing and managing services, Qdrant adds operational overhead. The managed cloud option helps but comes at a cost.
When to choose it: You need expressive filtering, you're comfortable with container-based deployments, and you want more control than a fully managed service offers.
Weaviate
Weaviate has a unique positioning: it tries to be both a vector database and a semantic knowledge graph, with built-in text vectorization and a GraphQL query interface.
I found the GraphQL API more complex than necessary for straightforward RAG. The built-in vectorizers are convenient if you want Weaviate to manage embedding generation, but they couple your embedding model choice to your database choice in ways that can be limiting.
When to choose it: You want built-in vectorization and are building something that benefits from Weaviate's knowledge graph features. For straightforward RAG, simpler options usually win.
What Actually Matters in Production
After building and operating this system for months, here's what I'd tell myself at the start:
Operational fit beats benchmark numbers. The system that fits your existing infrastructure and team's operational skills will outperform a "faster" system that introduces friction.
Metadata filtering design is critical. I spent significant time rethinking my metadata schema after the first version made common queries harder than they needed to be. Think carefully about what you need to filter on before you design your schema — the vector search layer should work with your filtering strategy, not against it.
Embedding model choice matters more than database choice. I tested several embedding models before settling on one for the support-automation pipeline. The quality of semantic search depends far more on how well the embedding model represents your domain than on which ANN algorithm the database uses. For support documentation, domain-specific fine-tuning of embeddings made a larger difference than any database tuning.
Re-embedding costs are real. When you change embedding models (and you will at some point), you need to re-embed your entire corpus. Design for this: keep your original text, track which embedding model version was used, and build a re-embedding pipeline before you need it.
Hybrid Search: The Gap Between Theory and Practice
Most production RAG systems end up doing some form of hybrid search — combining vector similarity with keyword matching (BM25) or metadata filters.
pgvector doesn't have built-in BM25, so in the support-automation system I handle this with a simple reranking step: retrieve more candidates from vector search, then apply keyword scoring before returning the final set.
This is a simple approach, but it works reliably. More sophisticated hybrid search — with proper BM25 indexing integrated into the vector store — is available in Weaviate and Qdrant out of the box.
My Recommendation by Context
There's no universally "best" vector database. Here's how I'd approach the choice:
You're adding vector search to an existing Postgres system: Start with pgvector. You get SQL flexibility, transactional consistency, and operational simplicity. Scale to a purpose-built solution later if you need to.
You're building a new service and want to ship fast without managing infrastructure: Pinecone or Qdrant Cloud. Accept the managed service tradeoff to move quickly.
You need expressive metadata filtering and are comfortable with containers: Qdrant self-hosted. The filter API is excellent and the Rust implementation handles concurrent load well.
Your collection is small (under a million vectors) and you need SQL joins: pgvector, no contest.
The support-automation RAG pipeline runs on pgvector. It's not the fastest possible choice, but it's the right fit: it lives alongside the relational data it searches against, it fails in ways I already know how to debug, and adding vector capabilities to an existing Postgres deployment took less than a day.
Start with the question "what fits my existing system?" rather than "what benchmarks best?" You'll ship faster and operate with less friction.