Skip to main content

How to use EvalGate

Welcome to the EvalGate help center. These guides explain, in plain English and with screenshots, how to use every part of EvalGate. No prior experience with evaluation platforms is required. Each article is a short, numbered walkthrough you can follow top to bottom.
Screenshots in this section show the EvalGate web app. The exact buttons and labels may shift between releases, but the steps stay the same. If a screen looks different, follow the words, not the picture.

New here? Start with EvalGate 101

Everything you need to understand the product in 15 minutes.

What EvalGate is for

One paragraph and one diagram: trace to eval to gate.

Finding your way around

The sidebar, the main panel, and where every feature lives.

Your first local gate

Block a regression in CI before lunch. No account required.

Connect a model provider

Bring your own key so model-backed features can run.

Evaluations & datasets

Build the test cases that define what “good” means for your AI.

Create an evaluation

From blank page to a runnable eval in two minutes.

Add test cases

Type, paste, import, or promote a real failure into a case.

Run an evaluation

One click, then read the pass/fail report.

Review the results

What the scores, costs, and failures actually mean.

Install an evaluation pack

Start with detailed healthcare, legal, support, coding, or financial coverage.

Repository intelligence

Find source-backed AI systems and missing coverage at one exact commit.

Scan a repository

Detect providers, agents, prompts, tools, evals, and release risks without executing repository code.

Traces & logs

See exactly what your AI did in production, then turn it into coverage.

Explore the logs

Filter thousands of traces down to the one that matters.

Inspect a trace

Tree, timeline, tools, cost, scores, and evidence.

Save views and cohorts

Keep a query so you never rebuild it twice.

Export traces

Download a bounded, redacted CSV or JSONL file.

Prompts & scorers

Version the prompts and scoring logic that decide quality.

Create a prompt

Author, version, and test a prompt before it ships.

Publish and roll back

Promote an approved version, then undo safely.

Create a scorer

Deterministic, code, composite, or LLM judge.

Calibrate a scorer

Anchor sets, mapping, and the release gate.

Playground & experiments

Safely try a change before it reaches users.

Run a Playground

Compare prompt variants side by side on real cases.

Pick a winner

Turn the best variant into the new baseline.

Continuous eval & monitors

Watch live traffic and get alerted when quality slips.

Create a monitor

Define what to sample and how to score it.

Simulate, then activate

Prove the monitor on history before it touches live traffic.

Read an alert

What a fired alert tells you and where to go next.

Copilot & review

Let the AI assistant draft changes, then keep a human in charge.

Ask the Copilot

Start a thread from any supported surface.

Accept or reject a proposal

Review the diff, then apply or throw it away.

Deployments & reports

Ship the winning configuration and seal the evidence.

Deploy an artifact

Promote a versioned, immutable artifact to production.

Run the eval loop

Carry one failure from trace to signed report.

Share a signed report

Give reviewers a verifiable, revocable summary.

Settings & account

Manage your workspace, team, keys, and billing.

Invite your team

Add members and set their permissions.

Manage API keys

Create, rotate, and revoke programmatic access.

Check feature status

See what is Beta, Experimental, or generally available.

Can’t find what you need?

Feature status

The authoritative inventory of what ships today.

Current-behavior manuals

Permissions, limits, failure behavior, and troubleshooting.

API reference

Integrate directly with the platform.

Open an issue

Report a bug or request a guide.