> ## Documentation Index
> Fetch the complete documentation index at: https://evalgate.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Measure retrieval with explicit reference evidence

> Use corpus and reference identities, meaningful denominators, and unavailable states when evidence is incomplete.

Retrieval quality concerns which evidence was returned and where it appeared. It is separate from whether a generated answer faithfully used that evidence. Use this guide alongside the broader [RAG evaluation guide](/docs/guides/rag-evaluation).

## Pin the protocol

The current built-in retrieval workflow is explicitly lexical. Its schema includes `method: "lexical"`, a bounded `topK`, a historical cutoff, an organization-owned document ID set, and a reference version. The metric computer is more general than this particular built-in execution route; do not infer a Jev reranker or arbitrary vector-database integration from the metric names.

The workflow pins corpus content and revalidates access and content before execution. A changed corpus or revoked access must not silently become comparable evidence from the original corpus.

## Label relevance explicitly

Each observation identifies its query, corpus, cutoff, retrieval unit, k, ordered returned IDs, relevance labels, and whether the complete relevant population is known.

The current reference gains are integers from 0 through 4. A zero gain explicitly means non-relevant. An absent label means unknown, not non-relevant.

Do not claim exhaustive recall against a corpus when only a partial set of relevant documents has been labeled.

## Metric semantics

| Metric | Interpretation in the current computer |
| - | - |
| Precision\@k | Relevant returned unique units divided by k |
| Recall\@k | Relevant returned units divided by the complete known relevant population; otherwise unavailable |
| nDCG\@k | Discounted graded relevance normalized against the ideal ranking from a complete reference population |
| MRR | Mean reciprocal rank of the first relevant result, subject to the documented missing-reference rules |

MRR is not the average rank position. A first relevant document at rank 4 contributes 1/4. A higher reciprocal-rank score means the first relevant result appeared earlier.

For example, with k = 3, returned IDs `[A, X, B]`, complete relevant set `[A, B, C]`, and explicit label X = 0: precision is 2/3, recall is 2/3, and reciprocal rank for this query is 1.

## Missing evidence is not zero quality

A partial scan, permission violation, missing reference set, or returned item without a label prevents the corresponding valid measurement. Duplicate reference identities are invalid. Returned duplicates are handled using first unique relevance units.

The aggregation does not quietly drop invalid queries and publish a selective-cohort mean. When a required query metric is unavailable, the aggregate remains unavailable with the missing count and reasons.

## Be precise about ranking ties

The current observation contract records ordered IDs, not a complete score distribution and tie-breaking provenance. It cannot, on its own, establish how much ranking quality came from a candidate scorer versus the original order used to resolve ties.

For an external reranking experiment, retain the raw candidate scores and tie policy in a separately reviewed experiment protocol. Report model-only and hybrid-system results separately; do not describe that diagnostic as an already shipped automatic EvalGate report.

Continue with [Compare evaluators](/docs/guides/compare-evaluators) and [Read evaluation evidence](/docs/guides/read-evaluation-evidence).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.