Report / Overview
Grape vs Vector RAG
Same documents, same questions, same LLM, two ways to find the passages it answers from. Open a suite to step through every question and read both answers side by side. Last run 2026-09-30 17:04.
English
28 questions on 12 English Wikipedia articles (tea, octopus, Marie Curie, ...), 630 KB. Four are asked in Hindi, Spanish or French, and five are off-topic.
Gujarati questions, English documents
The same 12 English articles, every question asked in Gujarati. Lexical search cannot match across languages, so Grape's first LLM call rewrites the question in the documents' language. Vector RAG uses a large multilingual embedding model (multilingual-e5-large), which embeds both languages in one space.
Gujarati questions, English documents, default English embedder
The same Gujarati questions, but Vector RAG uses the English embedding model most tutorials start with (bge-small-en). Shown to make the point that a vector setup only reaches across languages if its embedder does.
Gujarati documents
Gujarati questions on 12 Gujarati Wikipedia articles (Ahmedabad, Gir, Narmada, ...). Vector RAG uses multilingual-e5-large.
Large context: 300 articles, 18 MB
300 long English Wikipedia articles, about 4.6 million tokens: far more than any model's context window, so the model only sees what retrieval hands it. Shows answer quality and the cost of getting the index ready.
How it works
Both answer the same questions with the same LLM and the same instruction: answer in 1-2 sentences from the given text, in the language of the question, or reply "Not in the documents." Answers are checked by keyword first: each expected fact must appear (common Gujarati spellings are listed as alternatives). When no keyword matches, an LLM judge (the same model) is asked whether the answer still states every expected fact, because a foreign name written in Gujarati has many spellings. The judge is used the same way for Grape and Vector RAG, and judged answers are marked "(judge)" on the run pages. Off-topic questions must be refused, checked by keyword only.
- Grape searches a trigram index with the question itself (BM25 per line, plus title, lead and proximity signals). For English questions the server then reranks the 20 best candidates with a small cross-encoder (ms-marco-MiniLM-L-6-v2, CPU) and sends the best 5 passages of up to 600 characters; a clearly off-topic English question is refused with no LLM call. When the match is weak (another language, a vocabulary gap), one small call sees only the file titles and rewrites the question in the documents' language, or says it is off-topic. The answer call may ask for one more search in other words.
- Vector RAG splits documents into 1000-character chunks, embeds them locally (fastembed, CPU), and sends the 5 nearest chunks. English sets use bge-small-en-v1.5; Gujarati sets use multilingual-e5-large.
Cost is the Claude API list price that the claude CLI reports per call. Index / embedding time is measured once on an 8-core laptop CPU; the vectors are cached for later runs. LLM answers vary between runs, so each suite is run several times and the page shows the mean with the range. With 23-29 questions a suite, one question is 3-4 points: read small gaps as ties.
Limits. Keyword scoring can miss a correct answer written another way; every answer is on the suite pages so you can check. The question sets are small and written by us from the articles. Grape's ranking was tuned on the English and Gujarati-question sets; the Gujarati-documents and large sets were only run after tuning.
Documents: Wikipedia articles, CC BY-SA 4.0 (corpus/ in the repository). Run it yourself: see the README.