Skip to main content
Retrieval quality concerns which evidence was returned and where it appeared. It is separate from whether a generated answer faithfully used that evidence. Use this guide alongside the broader RAG evaluation guide.

Pin the protocol

The current built-in retrieval workflow is explicitly lexical. Its schema includes method: "lexical", a bounded topK, a historical cutoff, an organization-owned document ID set, and a reference version. The metric computer is more general than this particular built-in execution route; do not infer a Jev reranker or arbitrary vector-database integration from the metric names. The workflow pins corpus content and revalidates access and content before execution. A changed corpus or revoked access must not silently become comparable evidence from the original corpus.

Label relevance explicitly

Each observation identifies its query, corpus, cutoff, retrieval unit, k, ordered returned IDs, relevance labels, and whether the complete relevant population is known. The current reference gains are integers from 0 through 4. A zero gain explicitly means non-relevant. An absent label means unknown, not non-relevant. Do not claim exhaustive recall against a corpus when only a partial set of relevant documents has been labeled.

Metric semantics

MRR is not the average rank position. A first relevant document at rank 4 contributes 1/4. A higher reciprocal-rank score means the first relevant result appeared earlier. For example, with k = 3, returned IDs [A, X, B], complete relevant set [A, B, C], and explicit label X = 0: precision is 2/3, recall is 2/3, and reciprocal rank for this query is 1.

Missing evidence is not zero quality

A partial scan, permission violation, missing reference set, or returned item without a label prevents the corresponding valid measurement. Duplicate reference identities are invalid. Returned duplicates are handled using first unique relevance units. The aggregation does not quietly drop invalid queries and publish a selective-cohort mean. When a required query metric is unavailable, the aggregate remains unavailable with the missing count and reasons.

Be precise about ranking ties

The current observation contract records ordered IDs, not a complete score distribution and tie-breaking provenance. It cannot, on its own, establish how much ranking quality came from a candidate scorer versus the original order used to resolve ties. For an external reranking experiment, retain the raw candidate scores and tie policy in a separately reviewed experiment protocol. Report model-only and hybrid-system results separately; do not describe that diagnostic as an already shipped automatic EvalGate report. Continue with Compare evaluators and Read evaluation evidence.