Back to the Catalog
rag
retrieval
embeddings
evaluation
llm

RAG Without the Hand-Waving

24 questions

A practical quiz about how RAG systems split, retrieve, rank, cite, and evaluate information, and how they fail. Every explanation links to supporting research or technical documentation.

Questions

  1. Not answered. What two kinds of memory did the original RAG model combine?
  2. Not answered. Which statements describe the original RAG design?
  3. Not answered. A chunk splits a refund rule from its exception. What should you fix first?
  4. Not answered. Why can whole-document chunks work poorly for clause-level questions?
  5. Not answered. What can fixed-window chunk overlap change?
  6. Not answered. Why are dual encoders useful for first-stage retrieval over a large corpus?
  7. Not answered. For L2-normalized vectors, which score gives the same ranking as cosine similarity?
  8. Not answered. Dense retrieval misses an exact error code. What should you add?
  9. Not answered. What can go wrong with approximate nearest-neighbor search?
  10. Not answered. Three of four relevant chunks appear in the top five results. What is recall@5?
  11. Not answered. Why might increasing retrieval k from 5 to 50 make an answer worse?
  12. Not answered. With a rank constant of 60, what RRF score do ranks 1 and 4 produce?
  13. Not answered. Why can RRF combine BM25 and vector rankings without normalizing their raw scores?
  14. Not answered. Where should an expensive query-passage reranker run?
  15. Not answered. The retriever drops the only supporting passage. Can the reranker recover it?
  16. Not answered. What did the "Lost in the Middle" experiments find?
  17. Not answered. How many tokens remain in this context budget?
  18. Not answered. What must a citation evaluator check?
  19. Not answered. Why score provenance separately from answer accuracy?
  20. Not answered. The right passage was retrieved. What can still go wrong?
  21. Not answered. Which parts of a RAG pipeline should be evaluated separately?
  22. Not answered. Why is one embedding leaderboard score not enough to choose a model?
  23. Not answered. Retrieval returns an outdated policy after the source changes. What should you fix?
  24. Not answered. Which RAG security risks have been demonstrated?