Skip to main content
A judge configuration can be valid while its evidence is insufficient for a protected decision. EvalGate separates score comparability, task-specific qualification, and permission to act.

Calibration is not one universal approval

The Calibration control plane maintains reviewed anchor versions, score mappings, comparability, and drift. It answers whether scores can be interpreted on a compatible scale. A native decision judge also has calibration evidence requirements for release-blocking execution. A calibration identifier is not a certificate: the referenced evidence must resolve through the trusted execution path. Neither mechanism should be described as a blanket endorsement of the provider across all tasks.

Qualification is scoped to evidence and policy

The experiment result binds task, evaluator, reference set, observation set, sampling frame, and policy identities. Qualification evaluates the required metrics and their uncertainty, not just an attractive point estimate. In the current canonical builder, the principal qualification requirements are a failure-detection recall floor and an unsafe-auto-pass risk ceiling. A requirement may be demonstrated, violation, inconclusive, or insufficient_support. A small dataset with no observed unsafe passes can still have an upper confidence bound above the permitted risk. Zero observed events is not zero uncertainty. Minimum support is a policy input, not a universal promise that a fixed sample size is adequate for every task.

Know what the human approved

Independently verifiable reference truth has a separate contract for the verifier identity, version, content hash, and environment. Do not invent a reviewer to satisfy a record that should instead identify a deterministic verifier. When policy requires measured_result, protocol-only approval must not satisfy it. The schema supporting that requirement does not itself prove that every user-facing flow can obtain the stronger approval.

Qualification can pass while action authority remains absent

A frozen-policy simulation can produce useful evaluator evidence while its action gate remains abstain. The result explicitly separates qualification from activation authorization. Likewise, a release promotion must meet its own conditions. The inspected enforcement path checks canonical passing evidence, the authorized execution basis, active evaluator and completed validation bindings, the target environment, the current evidence identity, and the relevant action approval. This is why “qualified on a holdout” is not the same as “approved to deploy this artifact.”

Resolve a blocked state without weakening the policy

For missing references, obtain the required independent evidence. For changed identities, create or revalidate the applicable version. For inadequate support, run the planned confirmatory population rather than duplicating observations. For missing result approval, obtain the approval required by the policy rather than reusing a protocol signature. For native decision-judge calibration, the reviewed normal factory does not connect the optional lookup; that is an implementation dependency, not something fixed by entering a different version string. Continue with Troubleshoot inconclusive results and Read evaluation evidence. Current availability remains governed by Feature status.