Do RAG systems retrieve or memorise?
When a retrieval system answers correctly, is it reading the documents or reciting what the model memorised? A five-condition experiment that changes only the documents the model sees.
- Python
- LangChain
- ChromaDB
- Llama 3.1 8B / 3.3 70B
- DeBERTa NLI
- Groq
- GROBID
- Built
- 501 open-access papers parsed with GROBID into 34,502 chunks in ChromaDB with BGE embeddings. A five-condition experiment that varies only the supplied documents, to separate retrieval-grounded answers from parametric memory, across 120 cells with two Llama models (3.1 8B and 3.3 70B) through the Groq API.
- How I checked it
- The LLM-generated questions leaked their answers, so I rebuilt the evaluation set by hand and the conclusion reversed. I validated the LLM judge against two blind annotators (0.975 agreement, kappa 0.95), with an NLI model from a different family as an independent check.
- Result
- Retrieval drives correctness, not memory: accuracy rose from 33% with no documents to 92% with retrieval (McNemar p = 0.0001, replicated across both model sizes). Giving the model the exact source paper showed no detectable further gain (p = 0.5). Both models adopted a fluent counterfactual in 24 of 24 tests. Limits: 12 questions, one topic, one model family.
Giving the model the exact source paper showed no detectable further gain over standard retrieval (p = 0.5).
