Suites / Large context: 300 articles, 18 MB
Large context: 300 articles, 18 MB
300 long English Wikipedia articles, about 4.6 million tokens: far more than any model's context window, so the model only sees what retrieval hands it. Shows answer quality and the cost of getting the index ready.
3 runs / 300 files, 18.4 MB / model haiku, thinking off / Vector RAG embedder BAAI/bge-small-en-v1.5. Raw results: 2026-09-30 16:57 2026-09-30 17:01 2026-09-30 17:04
Grape
22.3 (22 to 23) / 23
- Off-topic refused
- 5.3 (5 to 6) / 6
- Input tokens per question
- 1,413 (1,386 to 1,431)
- LLM calls per question
- 1.02 (1.00 to 1.03)
- Cost per question
- $0.00167 ($0.00163 to $0.00169)
- Time per question
- 3.5 s (3.4 s to 3.7 s)
- Index time
- 1.1 s
Vector RAG
23 / 23
- Off-topic refused
- 6 / 6
- Input tokens per question
- 1,391
- LLM calls per question
- 1.00
- Cost per question
- $0.00163 ($0.00162 to $0.00164)
- Time per question
- 3.0 s (3.0 s to 3.0 s)
- Embedding time
- 46.3 min
Questions
right every run some runs never. Left half Grape, right half Vector RAG.
j / k or arrow keys