Grape benchmark

Suites / Large context: 300 articles, 18 MB

Large context: 300 articles, 18 MB

300 long English Wikipedia articles, about 4.6 million tokens: far more than any model's context window, so the model only sees what retrieval hands it. Shows answer quality and the cost of getting the index ready.

3 runs / 300 files, 18.4 MB / model haiku, thinking off / Vector RAG embedder BAAI/bge-small-en-v1.5. Raw results: 2026-09-30 16:57 2026-09-30 17:01 2026-09-30 17:04

Grape

22.3 (22 to 23) / 23

Off-topic refused
5.3 (5 to 6) / 6
Input tokens per question
1,413 (1,386 to 1,431)
LLM calls per question
1.02 (1.00 to 1.03)
Cost per question
$0.00167 ($0.00163 to $0.00169)
Time per question
3.5 s (3.4 s to 3.7 s)
Index time
1.1 s
Vector RAG

23 / 23

Off-topic refused
6 / 6
Input tokens per question
1,391
LLM calls per question
1.00
Cost per question
$0.00163 ($0.00162 to $0.00164)
Time per question
3.0 s (3.0 s to 3.0 s)
Embedding time
46.3 min
Questions

right every run some runs never. Left half Grape, right half Vector RAG.

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 15
  16. 16
  17. 17
  18. 18
  19. 19
  20. 20
  21. 21
  22. 22
  23. 23
  24. 24
  25. 25
  26. 26
  27. 27
  28. 28
  29. 29

j / k or arrow keys