23 May 2026 · 6 min
Teams obsess over the vector database and default the embedding model. It is exactly backwards.
When a team sets out to build semantic search, the conversation almost always starts in the wrong place. Weeks go into choosing a vector database — Qdrant or pgvector, this index or that, self-hosted or managed — and minutes go into the embedding model, which usually defaults to whatever is popular or whatever the tutorial used. This is exactly backwards. The database decides how fast and how cheaply you can search. The embedding model decides whether the right results come back at all.
If your search is returning irrelevant results, the odds are overwhelming that the problem is the embedding, not the database. And no amount of tuning the database will fix a bad embedding, because the mistake was made before the database ever saw the data.
An embedding model turns each piece of text into a point in a high-dimensional space, positioned so that similar meanings sit close together. Search then works by finding the points nearest to your query. This means the embedding is what defines "near." If the model places a query far from its true answer in that space, no database — however fast, however well-indexed — can retrieve it, because as far as the database is concerned they simply are not close.
Relevance, in other words, is determined at embedding time, not query time. The database is a very fast way of finding nearby points; it has no opinion about whether the points were placed sensibly in the first place. That placement is entirely the embedding model's job.
There is no single best embedding model, because "best" depends on what you are embedding. A model that excels at legal text may be mediocre on short product titles; one tuned for conversational support tickets may misjudge dense technical documentation. Domain, document length, vocabulary and structure all change which model performs. A model sitting at the top of a general leaderboard can be beaten, on your specific data, by a smaller model that happens to fit your content better.
This is why the choice cannot be made in the abstract. The question is never "what is the best embedding model" but "what is the best embedding model for this content and these queries" — and those are very different questions.
For anything beyond English, the model choice becomes even more consequential. A model trained predominantly on English will quietly underperform on Finnish, Norwegian, Swedish or Danish — not with an obvious error, but with subtly worse placements that surface as slightly-wrong search results. For multilingual content, or content where a query in one language should match a document in another, a genuinely multilingual embedding model is not a nice-to-have; it is the difference between search that works and search that almost works.
This is a common and costly oversight. A team benchmarks in English, everything looks fine, and the degradation only appears once real Nordic-language content and queries arrive — by which point the architecture is committed.
Public embedding leaderboards are useful as a starting shortlist and dangerous as a decision. They rank models on standardised benchmark datasets that are almost certainly not your data, in tasks that are almost certainly not your task. A model can top the leaderboard on generic retrieval and still underperform on your domain, your document lengths and your languages. Treating the leaderboard as a proxy for your use case is a bet that your data looks like the benchmark's — a bet that is usually wrong.
The leaderboard tells you which models are worth testing. It does not tell you which model to use. Only your data can do that.
The reliable way to choose is unglamorous: take a handful of candidate models, embed a representative slice of your real content, run your real queries against each, and measure which returns the right results. This does not take long, and it replaces guesswork with evidence. We do exactly this at the start of a search project, because an hour spent benchmarking on real data saves weeks of tuning a database to compensate for an embedding that was never going to work.
The measurement is the whole point. You are not looking for the model with the best reputation; you are looking for the one that places your queries closest to your answers, and that can only be found by trying them on your data.
How you split documents into chunks interacts directly with the embedding. Chunks that are too long dilute the meaning into a vague average; chunks that are too short lose the context that makes them findable. The right chunking strategy depends on the model and the content together, and getting it wrong can make even a good embedding model return poor results. It is not a separate decision made afterwards; it is part of the same design problem.
The embedding model and the chunking strategy are the foundation of semantic search. Get them right, benchmarked on your real data, and the database becomes what it should be: an implementation detail about speed and cost, chosen for operational reasons. Get them wrong, and no database will save you — you will spend your time tuning infrastructure to compensate for a foundation that was flawed from the start. Spend the effort where it actually determines the outcome, and the rest of the system gets much easier.
Book a 30-minute call. We will tell you honestly whether we can help.
Book a call