Skip to main content

Understand the product before writing evals

You should not have to explain tracing, install an OpenTelemetry collector, or fill out a long questionnaire to get a useful first eval. EvalGate starts by reading the product surface and teaching its working theory back to you. You correct that theory once, then use real examples to define what good behavior means.
This understand-and-author workflow is experimental. The TypeScript and Python CLIs are distributed through their package registries. Live provider-session capture and hosted delivery still require an explicit metrics export; EvalGate does not read private transcripts or run an always-on collector.

Install the canonical EvalGate Skill collection

Browse the official EvalGate Skills repository for source, references, and MCP configuration. Install the portable public collection:
Inspect the six-Skill collection before installing it:
The complete project-local collection is the supported public installation path. Selective installation is not advertised because the full installed-output reference graph is the governed portable contract. The collection is Experimental. Fixture and scorer tests prove contract behavior; they do not prove that a live model follows the Skill. The primary Skill decides whether an AI evaluation is needed, selects the smallest defensible coverage, distinguishes regressions from inconclusive or infrastructure-failed runs, and prevents threshold or baseline gaming. Five focused sibling Skills cover setup, regression gates, bounded traces, repository questions, and MCP. EvalGate does not expose its private application repository, customer data, prompts, or proprietary scoring logic through this distribution. When the agent returns a machine-readable decision, it uses one of nine canonical classifications and derives invokeEvalGate and releaseDecision from that classification. Read the published JSON Schema instead of inventing labels such as pass, fail, or needs_evals. MCP run, baseline, and project tools can supply evidence for the decision, but remain read-only and do not create a separate enum or authorize promotion. Every machine-readable decision also reports quality, protected slices, reliability, latency, and cost through the schema’s required evidenceSummary. A dimension without valid evidence must use status: "not_measured" and explain why; agents must not infer or fabricate a number merely to fill the report. Then ask:
Add evals to this product. First understand what it does from repository evidence, show me your working theory and one important question, and do not apply or promote anything without asking.
The skill calls the same EvalGate CLI and schemas used by local development and CI. It does not maintain a separate eval format and is not the enforcement boundary: the EvalGate runtime, reviewed baseline, policies, and CI gate supply execution and durable evidence. Teams can commit the installed Skill to share one reviewed version through source control. To evaluate a change to the Skill itself, compare an agent with no Skill, the reviewed Skill, and the candidate Skill as separate experiment variants. Preserve model/agent identity, repeated trials, task outcomes, tool behavior, tokens, cost, latency, and protected-slice evidence before promotion.

Connect the coding agent you already use

Run this from the repository you want the agent to understand. It previews the files first, then writes only after --apply:
Choose claude-code, cursor, cowork, or generic for another client. The installer uses the client’s repository instruction format where one is known:
  • Codex: .agents/skills/...
  • Claude Code: .claude/skills/...
  • Cursor: .cursor/rules/...mdc
  • CoWork and other clients: the portable .agents/skills/... layout
It also records the selected client and its metrics-only export path in .evalgate/agent-clients.json. Complete signed-in setup and configure the signed-in repository setup before running the installer; it does not mint an implicit activation key or start a collector. After the client exports bounded outcome and efficiency totals, capture them explicitly:
This is a repository installer, not an app-store registration. Publishing a new SDK package is required before a public npx install can use a newly added CLI command.

Get a useful working theory first

The agent begins with a read-only preview:
It should respond as soon as there is enough evidence for a useful first read. It should not make you wait for a comprehensive scan. The response should look roughly like this:
Working theory: This is a support assistant for customers asking about orders and refunds. The visible output is a policy-grounded answer with a concrete next step. Observed: The refund route, support prompt, and fixture data all reference a 30-day standard window. Hypotheses: The suite should probably protect policy accuracy, directness, and escalation when an exception is possible. Confidence: Medium. The core flow is clear, but no approved response style or exception behavior is saved yet. One question: When completeness and speed conflict, which should the customer notice first?
The labels matter:
  • Observed means repository content, tests, fixtures, saved behavior, or an example you supplied supports the statement.
  • Hypothesis means the agent inferred product intent, a failure mode, or a preferred response that you have not confirmed.
A plausible failure is not a real failure. An agent-written answer is not trusted product truth just because it sounds good. The preview performs no file writes, network calls, model calls, Git changes, baseline changes, or promotions.

Confirm one product context

Answer the single question that would most change the evals, then review the proposed context. After explicit approval, save it:
EvalGate writes the reviewed context to evalgate.product.json. That file is the canonical statement of product intent, evidence, observations, hypotheses, confidence, and open questions. Claude, Cursor, Codex, Copilot, and other native agent instruction files should only point to it:
Do not copy the product description or quality rules into every agent file. One canonical artifact prevents stale and contradictory instructions.

Describe quality with real contrasts

After confirming or correcting the product context, give the agent:
  • a realistic user input;
  • the preferred response;
  • a response to avoid;
  • what is observably wrong with the avoided response;
  • any product-specific priority or behavior that must never occur.
Use specific problems such as “invented refund eligibility” or “ignored the requested output format.” Avoid broad labels such as “bad quality.” If the agent proposes a preferred response, failure mode, or behavior to protect, it must remain labeled as a hypothesis until you edit or approve it. The agent proposes a versioned evalgate.quality.json:
The profile is editable repository data. It does not contain a hidden prompt or provider-specific judge configuration. Top-level dimensions are a coverage vocabulary. An example’s dimensionValues records only the scenario values that its input and outputs actually establish. Profile synthesis creates one draft per contrastive example; it does not manufacture a Cartesian set by relabeling an unchanged example. Add another real contrastive example when you want coverage for another value.

Preview before writing

Preview reports:
  • the normalized profile hash;
  • how many contrastive examples were supplied;
  • which quarantined cases would be created;
  • whether the output is new or would be replaced;
  • the exact next command.
Profile preview performs no file writes, network calls, model calls, Git mutations, baseline changes, or promotions.

Create quarantined drafts

After reviewing the profile and preview:
The command writes .evalgate/golden/synthetic.jsonl. Every row has lifecycleState: "synthetic" and records which profile and example produced it. Applying drafts does not make them a merge blocker. Review generated cases separately before running local promotion. Baseline acceptance and Git changes are also separate actions.

Add real usage when it helps

Real usage is an optional next source of examples, not a setup requirement. EvalGate uses these technical terms: Connect automatic capture when you need production discovery, latency and cost details, or multi-step debugging. Until then, pasted interactions, saved run files, API responses, and hand-written contrastive examples can all begin the evaluation loop. When you have representative real interactions:
  1. Label each pass or fail and note the first observable problem.
  2. Let product-specific failure categories emerge from the examples.
  3. Fix the underlying product behavior when possible.
  4. Create one binary evaluator per recurring subjective problem.
  5. Validate model-based judges against independent human labels before gating.
Synthetic cases fill coverage gaps. They do not replace human review or prove that production behavior is covered.