AmitSingh
All posts
AI / LLMs RAG LLM Architecture

Vectorless RAG: when you do not need embeddings

Vector databases became the default answer for retrieval. For a lot of enterprise knowledge work, they are more machinery than the problem deserves.

Amit Shivpratap Singh • • 6 min read

Reviewing retrieval metrics on a dashboard
On this page

The default that stopped being questioned

Ask how to ground a model in company knowledge and the answer arrives fully formed: chunk the documents, embed the chunks, store the vectors, retrieve by cosine similarity. It is a good pattern. It is not the only one, and it is rarely the cheapest one to operate.

Every embedding pipeline adds moving parts — a chunking strategy, an embedding model whose version you must pin, a vector store to host and back up, and a re-indexing job that runs whenever source content changes. Each part is a place where answers quietly go stale.

What vectorless retrieval actually means

Vectorless retrieval leans on structure that already exists. Enterprise content is rarely an undifferentiated pile of text: it lives in SharePoint libraries with metadata, in Dataverse tables with relationships, in ticketing systems with statuses and owners. That structure is retrieval signal you get for free.

Instead of embedding everything, you let the model plan a query — filter by product line, date range and document type — then hand it the handful of records that survive. Keyword search, SQL predicates and metadata filters do the narrowing; the model does the reasoning.

typescript
interface RetrievalPlan {
  productLine: string;
  documentType: "spec" | "policy" | "runbook";
  updatedSince: string;
}

// The LLM emits the plan; SQL enforces permissions and freshness.
async function retrieve(plan: RetrievalPlan) {
  return db.documents.findMany({
    where: {
      productLine: plan.productLine,
      type: plan.documentType,
      updatedAt: { gte: plan.updatedSince },
    },
    take: 12,
  });
}
The model plans a filter; the database does the narrowing.

Where it wins

It wins when the corpus is small enough to filter deterministically, when freshness matters more than recall, and when you need to explain exactly why a document was shown. A filter is auditable in a way a similarity score is not.

It also wins on cost. No embedding calls, no vector store, no re-indexing pipeline — and permissions come along for the ride, because you are querying the system that already enforces them.

Where it does not

Semantic search over millions of unstructured documents is still a vector problem. So is any case where users phrase questions in language that shares no vocabulary with the source material. The honest answer is usually hybrid: metadata filters to narrow the field, embeddings to rank what is left.

The point is not that vectors are wrong. It is that retrieval is a design decision, and reaching for a vector database before you have looked at the structure you already own is skipping the cheap step.

Share LinkedIn X Email

Keep reading