> ## Documentation Index
> Fetch the complete documentation index at: https://evalgate.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# How the evaluation system fits together

> Separate the AI application, its evaluators, measurement evidence, qualification, and permission to act.

EvalGate answers several different questions. They should not collapse into one score:

| Layer | Question | Example |
| - | - | - |
| Application behavior | What did the AI system actually do? | A support assistant promised a refund. |
| Evaluation | Did that behavior meet a defined requirement? | Was the promise supported by the supplied policy? |
| Evaluator measurement | Is the checker reliable on this task? | Does it detect unsupported promises without mislabeling supported ones? |
| Qualification | Is the available evidence sufficient for a stated policy? | Did the independent holdout meet the required bounds? |
| Authorization | May this actor use this result for this action? | Is promotion approved for this exact environment and artifact? |

A successful request, a valid typed answer, a passing case, a qualified evaluator, and an authorized release are different outcomes.

## Start with the object being changed

When you change the application prompt, model, tools, or retrieval pipeline, compare application behavior while keeping the evaluator stable where possible. The [Experiments](/docs/platform/experiments) workflow documents that comparison.

When you change the evaluator, keep the application responses and admissible reference evidence fixed. Otherwise, a better response can be mistaken for a better evaluator. Use the [evaluator comparison guide](/docs/guides/compare-evaluators) for that distinction.

When you change an operating threshold without executing the corresponding real-world action, describe the result as a policy simulation. A simulated block is not evidence that a production action was prevented.

## Choose a checker that matches the requirement

Use a deterministic rule for a requirement that can be checked exactly: an allowed value, a schema, or an arithmetic invariant. Use a generative judge when a reviewed rubric requires an explanation or other generated text. A decision judge such as Jev returns bounded typed observations instead of a written critique.

The provider is one part of the evaluator. The questions, acceptance policy, input construction, model identity, and aggregation also affect what its result means.

## Understand the four similar-sounding records

**A scorer version** defines a reusable check. **An evaluator release** identifies a measurement instrument and its protocol. **A validation run** measures that instrument using selected evidence. **An application experiment** compares application variants. They can reference each other, but they are not interchangeable names for a run.

Likewise, the Calibration control plane's approved score mapping supports score comparability. It does not by itself demonstrate task-specific failure detection or grant deployment permission.

## Read uncertainty as information

An evaluator can return an observation without an accepted pass/fail interpretation. A run can execute completely while still being inconclusive. A comparison can be statistically promising but lack the independent evidence required for qualification.

Before acting, identify the original case, the final composed outcome, any unresolved attempts, the measurement population, and the policy being applied. Never treat `passed: false` alone as proof of a behavioral violation: it can also accompany non-authoritative outcomes.

## Availability boundary

EvalGate's feature inventory distinguishes controlled-beta workflows from experimental source capabilities and unverified release paths. This conceptual map does not change those labels. A feature appearing in the current source does not establish availability in every installed package or deployment.

Continue with [Decision judges](/docs/platform/decision-judges), [Qualification and authorization](/docs/platform/evaluator-qualification), and [Reading evaluation evidence](/docs/guides/read-evaluation-evidence).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.