The local loop (
auto contract, auto freeze, auto report, auto export, auto propose) is in the TypeScript CLI source. Before relying on it, check that your installed CLI lists those subcommands: npx @evalgate/sdk auto --help. The Python CLI does not implement this loop; from Python, use evalgate api for the hosted steps.The journey at a glance
You are asked to act only where a decision is yours: credentials, which production data may be used, how much may be spent, approving the plan, and adopting or merging the result.
1. Point EvalGate at the feature
Link the repository on/repository-intelligence and pick the AI feature the scan found: the prompt file, its entrypoint, and the model it calls. Bind it to a remote runner you control so EvalGate can execute the real feature, not a mock. Locally, npx @evalgate/sdk understand previews the same theory without credentials.
2. Gather cases and a grader
You need cases that show both what works and what fails, plus a grader you trust.- Start from failures you already have: labeled traces, review-queue decisions, or a dataset.
- Ask Copilot. It answers dataset questions from real queries and tells you exactly what it checked: “Searched in Support tickets · 3 of 4,000 rows matched · 3 read”, with the row ids it cited. It can propose a new draft dataset or evaluation that you review and accept.
- A failed production answer is an observation, not a case. A case needs a label or an approved grader.
- Development cases: the optimizer sees these failures and learns from them.
- Selection cases: candidates are measured on these, and any regression discards the candidate.
- Confirmation cases: never run, shown or read by the optimizer. They are kept back for the independent check in step 7.
3. Say what should improve
Write what you want the way you would tell a colleague:4. Approve cases, grader and spend
5. Try candidates
6. Freeze, report, export
.evalgate/auto/report.md (every candidate, pass rates with intervals, cost and latency with their boundaries, cheaper candidates that broke something, the exact sentences changed), writes .evalgate/auto/campaign.json, and opens a draft pull request. The report says the candidate is not confirmed yet.
7. Confirm on cases the optimizer never saw
Open/change-assessments/import and upload campaign.json. EvalGate binds the candidate commit, then runs the approved confirmation cases on the base and the candidate through your runners. The page shows:
- quality per protected metric with its margin, and safety failures;
- operating cost per request, the saving, and break-even, only when the runners reported usage on both sides;
- a claim: cost reduction supported, selection evidence only, refused, or no claim.
8. Start release review and merge
When confirmation holds, a person signed in to EvalGate starts release review from the authority panel. Agents, API keys and MCP tokens cannot take this step. You then review and merge the draft pull request in your code host. EvalGate never merges.Four examples
Support assistant: stop refunding before confirmation
Support assistant: stop refunding before confirmation
Cases: review-queue items where the assistant issued a refund without confirming, plus ordinary refund conversations that went well.Intent: “Fix the refund confirmation failure in the prompt. Keep accuracy within 2%. At most 5 experiments.”What protects you: a candidate that asks for confirmation but starts refusing legitimate refunds regresses selection cases and is discarded. The confirmation run checks the kept candidate on refund conversations it never saw.
RAG answerer: cite the source, stay as accurate
RAG answerer: cite the source, stay as accurate
Cases: questions with known answers and the documents that support them; label answers that cite nothing or the wrong document.Intent: “Make answers cite the retrieved document. Keep correctness within 1%. Spend up to $5.”What protects you: correctness is a protected metric with a 1% margin, so a candidate that cites more but answers worse is not kept. Retrieval code is not editable under the contract; only the prompt changes.
Lead qualification: cut cost without losing good leads
Lead qualification: cut cost without losing good leads
Cases: past inbound leads labeled qualified or not, including the expensive edge cases.Intent: “Cut cost by changing the prompt. Keep accuracy within 5%. Try cheaper models.”What protects you: a cost candidate is kept only if it is cheaper on measured cost and stays within the margin case by case. If either side’s cost is unknown, no saving is claimed. The confirmation run reports savings per request and, when the optimization spend was recorded, how many requests it takes to pay that spend back.
A feature in your own repository
A feature in your own repository
Cases: the eval specs already in the repo (
evals/), tagged with family:<name> so related cases stay together.Intent: “Fix the tone mismatch in prompts/support.md. Keep all existing specs passing.”What protects you: spec files, grader helpers and evals/baseline.json are immutable under the contract; editing them to pass is refused, not rewarded.What is not automated
- Approving a contract, starting release review, adopting a candidate, and merging always need a person.
- A scheduled campaign runs only what a person approved.
npx @evalgate/sdk auto dispatch --open-pr --state-remote originfrom a nightly CI job runs the approved plan in an isolated worktree and opens a draft pull request. It refuses when the plan changed since approval, when the approval is older than 14 days, when the base has uncommitted changes, when another dispatch is running, or when the approval has been used three times.--state-remote originkeeps those limits on the remote, because every CI run starts from a fresh checkout. See the CLI reference for the workflow file. - A paid confirmation run spends real model budget on your runners; set the budget in the contract and on the runner.