Vectorless RAG: when you do not need embeddings
Vector databases became the default answer for retrieval. For a lot of enterprise knowledge work, they are more machinery than the problem deserves.
Amit Shivpratap Singh • • 6 min read
On this page
The default that stopped being questioned
Ask how to ground a model in company knowledge and the answer arrives fully formed: chunk the documents, embed the chunks, store the vectors, retrieve by cosine similarity. It is a good pattern. It is not the only one, and it is rarely the cheapest one to operate.
Every embedding pipeline adds moving parts — a chunking strategy, an embedding model whose version you must pin, a vector store to host and back up, and a re-indexing job that runs whenever source content changes. Each part is a place where answers quietly go stale.
What vectorless retrieval actually means
Vectorless retrieval leans on structure that already exists. Enterprise content is rarely an undifferentiated pile of text: it lives in SharePoint libraries with metadata, in Dataverse tables with relationships, in ticketing systems with statuses and owners. That structure is retrieval signal you get for free.
Instead of embedding everything, you let the model plan a query — filter by product line, date range and document type — then hand it the handful of records that survive. Keyword search, SQL predicates and metadata filters do the narrowing; the model does the reasoning.
interface RetrievalPlan {
productLine: string;
documentType: "spec" | "policy" | "runbook";
updatedSince: string;
}
// The LLM emits the plan; SQL enforces permissions and freshness.
async function retrieve(plan: RetrievalPlan) {
return db.documents.findMany({
where: {
productLine: plan.productLine,
type: plan.documentType,
updatedAt: { gte: plan.updatedSince },
},
take: 12,
});
}
Where it wins
It wins when the corpus is small enough to filter deterministically, when freshness matters more than recall, and when you need to explain exactly why a document was shown. A filter is auditable in a way a similarity score is not.
It also wins on cost. No embedding calls, no vector store, no re-indexing pipeline — and permissions come along for the ride, because you are querying the system that already enforces them.
Where it does not
Semantic search over millions of unstructured documents is still a vector problem. So is any case where users phrase questions in language that shares no vocabulary with the source material. The honest answer is usually hybrid: metadata filters to narrow the field, embeddings to rank what is left.
The point is not that vectors are wrong. It is that retrieval is a design decision, and reaching for a vector database before you have looked at the structure you already own is skipping the cheap step.