Define what the comparison can establish
Record the task, error types, protected populations, source of reference labels, and decision you intend to make. State whether the candidate must improve failure detection, reduce cost within an agreed quality margin, or satisfy a particular absolute requirement. Keep three claims separate: “matches the baseline,” “matches independently verified references,” and “gives the same answer when repeated.” They measure agreement, correctness, and repeatability respectively.Freeze the comparable inputs
Pin the baseline validation run and its measurement revision. For an evaluator comparison, both arms should judge the same application input, response, evidence construction, reference labels, and environment. EvalGate’s persisted comparison resolves a pinned baseline revision and observation identity. The evaluator-comparison projection binds the target output into input identity, rather than treating two different responses as the same judging task. The comparison descriptor includesbaselineValidationRunId, purpose, taskId, minimumReleaseSupport, recallFloor, unsafeAutoPassCeiling, and an optional requiredApprovalStage. This is a description of fields, not a request to hide server-owned qualification or execution metadata inside a sampling manifest.