Grape benchmark

Report / Overview

Grape vs Vector RAG

Same documents, same questions, same LLM, two ways to find the passages it answers from. Open a suite to step through every question and read both answers side by side. Last run 2026-09-30 17:04.

How it works

Both answer the same questions with the same LLM and the same instruction: answer in 1-2 sentences from the given text, in the language of the question, or reply "Not in the documents." Answers are checked by keyword first: each expected fact must appear (common Gujarati spellings are listed as alternatives). When no keyword matches, an LLM judge (the same model) is asked whether the answer still states every expected fact, because a foreign name written in Gujarati has many spellings. The judge is used the same way for Grape and Vector RAG, and judged answers are marked "(judge)" on the run pages. Off-topic questions must be refused, checked by keyword only.

  • Grape searches a trigram index with the question itself (BM25 per line, plus title, lead and proximity signals). For English questions the server then reranks the 20 best candidates with a small cross-encoder (ms-marco-MiniLM-L-6-v2, CPU) and sends the best 5 passages of up to 600 characters; a clearly off-topic English question is refused with no LLM call. When the match is weak (another language, a vocabulary gap), one small call sees only the file titles and rewrites the question in the documents' language, or says it is off-topic. The answer call may ask for one more search in other words.
  • Vector RAG splits documents into 1000-character chunks, embeds them locally (fastembed, CPU), and sends the 5 nearest chunks. English sets use bge-small-en-v1.5; Gujarati sets use multilingual-e5-large.

Cost is the Claude API list price that the claude CLI reports per call. Index / embedding time is measured once on an 8-core laptop CPU; the vectors are cached for later runs. LLM answers vary between runs, so each suite is run several times and the page shows the mean with the range. With 23-29 questions a suite, one question is 3-4 points: read small gaps as ties.

Limits. Keyword scoring can miss a correct answer written another way; every answer is on the suite pages so you can check. The question sets are small and written by us from the articles. Grape's ranking was tuned on the English and Gujarati-question sets; the Gujarati-documents and large sets were only run after tuning.

Documents: Wikipedia articles, CC BY-SA 4.0 (corpus/ in the repository). Run it yourself: see the README.