Skip to main content
This guide walks through one full improvement campaign: you say in plain words what should get better, EvalGate tries candidate prompt edits on an isolated branch under limits you approve, confirms the best candidate on cases the optimizer never saw, and hands you a report and a draft pull request. Nothing merges or deploys unless a person does it.
The local loop (auto contract, auto freeze, auto report, auto export, auto propose) is in the TypeScript CLI source. Before relying on it, check that your installed CLI lists those subcommands: npx @evalgate/sdk auto --help. The Python CLI does not implement this loop; from Python, use evalgate api for the hosted steps.

The journey at a glance

You are asked to act only where a decision is yours: credentials, which production data may be used, how much may be spent, approving the plan, and adopting or merging the result.

1. Point EvalGate at the feature

Link the repository on /repository-intelligence and pick the AI feature the scan found: the prompt file, its entrypoint, and the model it calls. Bind it to a remote runner you control so EvalGate can execute the real feature, not a mock. Locally, npx @evalgate/sdk understand previews the same theory without credentials.

2. Gather cases and a grader

You need cases that show both what works and what fails, plus a grader you trust.
  • Start from failures you already have: labeled traces, review-queue decisions, or a dataset.
  • Ask Copilot. It answers dataset questions from real queries and tells you exactly what it checked: “Searched in Support tickets · 3 of 4,000 rows matched · 3 read”, with the row ids it cited. It can propose a new draft dataset or evaluation that you review and accept.
  • A failed production answer is an observation, not a case. A case needs a label or an approved grader.
EvalGate splits your cases into three groups by source, so near-duplicates never straddle them:
  • Development cases: the optimizer sees these failures and learns from them.
  • Selection cases: candidates are measured on these, and any regression discards the candidate.
  • Confirmation cases: never run, shown or read by the optimizer. They are kept back for the independent check in step 7.

3. Say what should improve

Write what you want the way you would tell a colleague:
EvalGate compiles this into an Optimization Contract: what may change (only the prompt), what is protected and by how much, how many candidates, how much spend, and that only a person adopts. Anything it did not understand is listed as Not understood rather than silently dropped; you either rephrase it or accept it explicitly. The same contract can be compiled and approved from the web, on an improvement cycle, before a run.

4. Approve cases, grader and spend

The review shows populations (confirmation cases only as counts and a digest), graders, protected slices, the mutation target, the spend bound and the contract, each with a risk level. High-risk items must be accepted by name. Changing any of it later requires a new approval.

5. Try candidates

The loop edits only the prompt. If a candidate’s run changes a spec, a grader, a dataset or the baseline, the candidate is refused and the files are restored. It stops at the trial limit and before spend can exceed the limit. Failed candidates are kept with the reason, so you can see what was tried.

6. Freeze, report, export

This pins the kept candidate by content hash on an isolated branch, writes .evalgate/auto/report.md (every candidate, pass rates with intervals, cost and latency with their boundaries, cheaper candidates that broke something, the exact sentences changed), writes .evalgate/auto/campaign.json, and opens a draft pull request. The report says the candidate is not confirmed yet.

7. Confirm on cases the optimizer never saw

Open /change-assessments/import and upload campaign.json. EvalGate binds the candidate commit, then runs the approved confirmation cases on the base and the candidate through your runners. The page shows:
  • quality per protected metric with its margin, and safety failures;
  • operating cost per request, the saving, and break-even, only when the runners reported usage on both sides;
  • a claim: cost reduction supported, selection evidence only, refused, or no claim.
A candidate that improved development cases but regressed confirmation cases gets no confirmed label and no adoption authority.

8. Start release review and merge

When confirmation holds, a person signed in to EvalGate starts release review from the authority panel. Agents, API keys and MCP tokens cannot take this step. You then review and merge the draft pull request in your code host. EvalGate never merges.

Four examples

Cases: review-queue items where the assistant issued a refund without confirming, plus ordinary refund conversations that went well.Intent: “Fix the refund confirmation failure in the prompt. Keep accuracy within 2%. At most 5 experiments.”What protects you: a candidate that asks for confirmation but starts refusing legitimate refunds regresses selection cases and is discarded. The confirmation run checks the kept candidate on refund conversations it never saw.
Cases: questions with known answers and the documents that support them; label answers that cite nothing or the wrong document.Intent: “Make answers cite the retrieved document. Keep correctness within 1%. Spend up to $5.”What protects you: correctness is a protected metric with a 1% margin, so a candidate that cites more but answers worse is not kept. Retrieval code is not editable under the contract; only the prompt changes.
Cases: past inbound leads labeled qualified or not, including the expensive edge cases.Intent: “Cut cost by changing the prompt. Keep accuracy within 5%. Try cheaper models.”What protects you: a cost candidate is kept only if it is cheaper on measured cost and stays within the margin case by case. If either side’s cost is unknown, no saving is claimed. The confirmation run reports savings per request and, when the optimization spend was recorded, how many requests it takes to pay that spend back.
Cases: the eval specs already in the repo (evals/), tagged with family:<name> so related cases stay together.Intent: “Fix the tone mismatch in prompts/support.md. Keep all existing specs passing.”What protects you: spec files, grader helpers and evals/baseline.json are immutable under the contract; editing them to pass is refused, not rewarded.

What is not automated

  • Approving a contract, starting release review, adopting a candidate, and merging always need a person.
  • A scheduled campaign runs only what a person approved. npx @evalgate/sdk auto dispatch --open-pr --state-remote origin from a nightly CI job runs the approved plan in an isolated worktree and opens a draft pull request. It refuses when the plan changed since approval, when the approval is older than 14 days, when the base has uncommitted changes, when another dispatch is running, or when the approval has been used three times. --state-remote origin keeps those limits on the remote, because every CI run starts from a fresh checkout. See the CLI reference for the workflow file.
  • A paid confirmation run spends real model budget on your runners; set the budget in the contract and on the runner.