24 June 2026 · 9 min
Vector search retrieves what looks similar. For multi-hop questions about relationships, that is not enough.
Retrieval-augmented generation has become the default way to ground a language model in your own data, and for most teams that means one thing: embed everything into vectors, store them in a vector database, and retrieve the nearest matches for each question. This works remarkably well for a large class of problems. But there is a second class where it quietly fails, and the failure is easy to miss because the system still returns confident, plausible answers — just incomplete or wrong ones.
The distinction that matters is not which technology is newer or more sophisticated. It is the shape of your data and the shape of the questions people ask of it. Get that diagnosis right and the choice between plain vector search and a graph-based approach becomes obvious.
Vector search answers questions of similarity. You have a large body of text, someone asks a question, and the answer lives in one passage — or a handful of passages — that are semantically close to the question. Product documentation, knowledge bases, support archives, research libraries: these are a near-perfect fit. The relevant fact sits in a chunk, the chunk is similar to the query, and retrieval finds it.
For this class of problem, vector search is fast, cheap, and hard to beat. Adding graph machinery on top would be over-engineering — complexity that buys nothing. If your questions are answered by finding the most relevant passage, you do not need anything more elaborate, and we will tell you so rather than sell you a bigger system.
The trouble begins when the answer does not live in any single passage but in the relationships between passages. Consider a question like: which suppliers are affected by a regulation that amends another regulation governing a component used in a product line owned by a particular subsidiary? No single chunk contains that answer. It has to be assembled by following links from one fact to the next.
Vector search cannot follow links. It retrieves passages that look similar to the question, but similarity is not connection. The individual hops in that supplier question may each be textually dissimilar to the original query, so the right passages never surface. The model then answers from whatever it did retrieve — fluently, confidently, and wrong. This is the failure mode that catches teams off guard, because nothing errors; the answer just quietly omits what it could not reach.
A knowledge graph stores facts as explicit entities and relationships: this regulation amends that one, this component belongs to that product, this product is owned by that subsidiary. GraphRAG combines this structure with retrieval. Instead of only finding similar text, it can traverse the graph — follow the amendment link, then the component link, then the ownership link — and assemble an answer from a chain of connected facts.
This unlocks multi-hop questions that flat retrieval cannot touch. It also brings a second benefit that matters enormously in regulated or high-stakes settings: traceability. Because the answer was built by walking an explicit path through the graph, you can show that path. The answer arrives with its reasoning attached, which reviewers can verify rather than trust.
Graphs are not free. Building one means extracting entities and relationships from your data — a modelling effort that plain vector search skips entirely. You have to decide what the entities are, what relationships connect them, and keep that structure current as the underlying data changes. For a body of text where the answers really do live in single passages, this is pure overhead with no return.
This is exactly why the diagnosis comes first. Reaching for a graph because it sounds more powerful is a common and expensive mistake. The graph earns its cost only when your questions genuinely require following connections; otherwise it is complexity you will maintain forever for no benefit.
In practice the strongest systems are rarely pure. Most real corpora contain both kinds of question: many that a single relevant passage answers, and some that require connecting facts. The right architecture uses vector search for semantic recall and a graph layer for the relational, multi-hop cases, routing each question to the mechanism that fits it.
Designing that split well — deciding what to model as a graph, what to leave as flat text, and how to combine the two at query time — is where most of the engineering judgement lives. Done right, you pay for graph complexity only where it pays you back, and vector search handles the rest cheaply.
You can usually diagnose your case from the questions alone. Write down the ten hardest questions your users actually ask. If each one is answered by finding the single most relevant passage, you need vector search and nothing more. If some of them can only be answered by combining several facts that are not textually similar to each other, you have multi-hop questions, and a graph layer will earn its place.
We start every retrieval project with exactly this exercise, on your real questions and your real data, before writing any code. The technology follows the diagnosis, never the other way around — and the diagnosis is almost always clearer than the marketing around either approach suggests.
Book a 30-minute call. We will tell you honestly whether we can help.
Book a call