Skip to main content
A decision judge evaluates supplied state using a fixed set of questions. EvalGate’s TypeSafe integration uses its native decision interface rather than asking a generative model to emit JSON. This is a guide to the current source contract. Check Feature status and your deployed contract before configuring it. It is not a claim that every decision-judge mode is production-qualified.

Three question types

The adapter validates the requested answer keys and probability distributions. Invalid evidence is rejected rather than repaired into a valid-looking observation. Valid output structure is not proof that the substantive answer is correct.

Acceptance gives an observation its meaning

A high probability is not automatically a good outcome. For “Does this answer contain an unsupported commitment?”, a high probability can mean strong evidence of a problem. Use choice_labels, score_threshold, or noul_polarity to declare the acceptance rule. A question without an acceptance policy remains observational rather than receiving an invented pass/fail interpretation. This is a question fragment, not a complete API request or a calibrated production policy:
The complete judge also needs its configured model, questions, aggregation, threshold, and execution context. Choose thresholds from task-specific evidence, not from an example in a guide.

Separate insufficient evidence from a negative judgment

A check should not call a promise unsupported merely because the required policy text was unavailable. Companion questions named in evidenceQuestionIds can require sufficient evidence before the primary outcome applies. Keep the primary check and evidence-sufficiency check separate. Missing or non-passing companion evidence produces insufficient_evidence on the shared disposition path; calling a more expensive model does not repair missing admissible evidence.

Preserve the whole result

Inspect decisions, decisionProvenance, child judge results, execution completeness, and the final composed outcome together. Provenance includes question identity, input construction, model identity, operating-policy identity, and any calibration reference. Native Jev results do not include a generated rationale. The stored marker is decision-model:rationale_unsupported. A UI can explain the configured acceptance rule without pretending that explanation is the model’s hidden reasoning.

Confidence and escalation

Confidence bands can retain a decision, request a frontier judge, or refer a case for human review. These are operating policies, not proof of empirical calibration. Native probability concentration must not be relabeled as “validated accuracy.” A successful fallback may resolve the final outcome while the original Jev observation still records an escalation request. Downstream reports must distinguish the original request from the terminal composed decision.

Credentials and release-blocking

Customer decision evaluations require the organization’s TypeSafe credential; the normal customer path does not silently borrow a platform environment key. Platform assistance is a separate, server-controlled purpose. Native release-blocking requires resolvable calibration evidence, not only a version string. In the reviewed source, the normal judge factory does not supply the optional calibration lookup, so this path fails closed. Do not document a settings-only production activation recipe until that resolver is connected and verified. Continue with Compare evaluators and Evaluator qualification.