Calibrate a scorer
Calibration makes a scorer’s scores comparable across versions and models. Without it, a “0.8” from one judge does not mean the same thing as a “0.8” from another. Calibration lives at/calibration.
Step 1 — Build an anchor set
An anchor set is a small, stable set of labeled cases that represent the range of quality you care about. Create one, label every item, and version it.
Step 2 — Run the scorer on the anchors
Run the current scorer version against the anchor set. EvalGate records the raw scores and the labels.Step 3 — Review the mapping
The mapping converts raw scores to a normalized scale (for example, 0–100) so different scorers can be compared. Review the mapping, then submit it for review.