Pay only for completed eval results
Your first 10,000 test-case results each month are free. After that, pay $1 per 1,000 while your model usage stays on your own provider bill.
First proof in 5 minutes
Start local, create a baseline, and fail a PR when behavior regresses.
Every miss becomes coverage
Promoted failures move from trace evidence into reusable eval checkpoints.
Reviewers get evidence
Gate reports, judge scores, and failure-mode deltas travel with the release.
One plan for every team
Pay as you go
Start free, then pay only for completed test-case results after the included allowance.
$1 per 1,000 eval results after 10,000/month
Get started freeNo card required for the free allowance.
Everything included
- 10,000 evaluation results free each month
- Unlimited projects, datasets, and runs
- Unlimited team members
- BYOK model usage with no EvalGate inference markup
- Passes and evaluated failures count once
- Retries are deduplicated by evaluation run
- $50 default monthly overage cap
Predictable examples
Monthly EvalGate charges, excluding your model provider bill.
10,000 results
$0
Included each month
50,000 results
$40
40,000 paid results
60,000 results
$50
Default monthly cap
What counts as an eval result?
The billable unit follows the durable evidence in your run—not model calls, scorer count, or empty run attempts.
One persisted result
A completed test case counts once, whether it passes or produces an evaluated failure.
Scorers do not multiply it
One case with several scorers is still one evaluation result.
Retries are deduplicated
A retry for the same evaluation run reuses its billing idempotency key.
Bring your own keys
OpenAI, Anthropic, and other model usage stays on your provider bill with no EvalGate inference markup.
Card only when needed
Start without a card. Add one only when you want to exceed 10,000 results in a month.
Spend is bounded
The default $50 monthly overage cap stops additional paid evaluation usage until it is raised.