| Playground | Beta | /playgrounds, then /evaluations/{id}/playground | get_playgrounds, get/post/patch_evaluations_id_playground, plus documented winner, variant, scorer, run, proposal, case, assistant, cancellation, and export operations | No dedicated client methods; use REST | API and persistence tests; exact-head CI | Evaluation Workflows |
| Prompt Hub | Beta | /prompts and an evaluation detail page | get_prompts and the evaluation prompt/version operation family | CLI sync supports prompt pull/push; no complete convenience-client surface | Prompt route tests; exact-head CI | Evaluation Workflows |
| Scorer Studio | Beta | /scorers and /scorers/{scorerId} | /api/scorers plus version, example, test, alignment, calibration, publish, binding, and usage operations | No dedicated SDK or CLI; use the authenticated REST workflow | Scorer workflow proof and hostile-code OCI boundary proof; exact-head CI | Evaluation Workflows |
| Red-Team Workbench | Beta | /red-team | /api/red-team risk-pack, campaign, run, finding, remediation, and signed-report operations | No dedicated SDK or CLI; use the authenticated REST workflow | Red-team workflow proof; exact-head CI | Evaluation Workflows |
| Logs & Trace Explorer | Beta | /logs | /api/logs query, detail continuation, aggregate, saved-view, cohort, bulk-action, and audited-export operations | No dedicated TypeScript, Python, or CLI convenience workflow; use authenticated REST | Unit, API, database, DOM, and golden-path proof source; exact-head CI | Production Intelligence |
| EvalGate Copilot | Beta | /copilot and supported product surfaces | /api/copilot thread, retained-context, message, action, proposal, accept/apply, and reject operations | No dedicated SDK or CLI convenience workflow; use authenticated REST and the review UI | Unit, API, database, DOM, and golden-path proof source; exact-head CI | Evaluation Workflows |
| Remote Runners | Beta | /remote-runners | /api/remote-runners registration, credentials, lifecycle, enqueue, claim, heartbeat, stream, cancellation, lease recovery, completion, and reconciliation operations | TypeScript and Python worker SDKs implement the signed protocol; no complete operator CLI | Unit, API, PostgreSQL, DOM, and golden-path proof source; exact-head CI | Evaluation Runtime |
| Deployable Assets | Beta | /deployable-assets | /api/deployable-assets list/create, approval, transition, rollout, health, and governed invocation operations | Generated contract types only; no complete TypeScript, Python, or CLI deployment client | Unit, API, PostgreSQL, DOM, and golden-path proof source; exact-head CI | Evaluation Runtime |
| Trajectory Analysis | Beta | /trajectory-analysis | /api/trajectory-analysis ingestion, detail, scoring, comparison, report, and retention operations | TypeScript and Python clients cover ingestion, inspection, scoring, reporting, and comparison; no dedicated CLI command | Unit, API, PostgreSQL, and DOM proof sources; isolated browser acceptance is tracked as pending; exact-head CI | Evaluation Quality |
| Dashboards and monitors | Beta | /dashboards and pinned /dashboards/shared/{token} views | Metric-definition, metric-query, dashboard/version/widget/render/share, and stateful alert/event/delivery operation families | No complete TypeScript, Python, or CLI convenience workflow; use the authenticated REST contract | Unit, API, PostgreSQL, DOM, and golden-path proof source; exact-head CI | Evidence & Reporting |
| Failure Topics | Beta | /insights | Production-insights summary, topic discovery/detail/version controls, merge/split, Gateway label proposals, and governed corrective actions | No dedicated TypeScript, Python, or CLI convenience workflow; use authenticated REST | Unit, API, PostgreSQL, DOM, and golden-path proof source; exact-head CI | Production Intelligence |
| Dataset Hub | Beta | Evaluation detail: test cases and labeled cases | get_evaluations_id_test_cases, post_evaluations_id_test_cases, get_evaluations_id_labeled_cases, post_evaluations_id_labeled_cases | TypeScript and Python workflows read canonical JSONL; CLI label, cluster, and sync commands | Golden lifecycle DB tests; exact-head CI | Evaluation Workflows |
| Experiments | Beta | /evaluations/{id} run and Auto panels | get_evaluations_id_runs, post_evaluations_id_runs, and the Auto-session operation family | CLI run, auto, and report artifacts; web Auto orchestration has no dedicated convenience client | Experiment runner tests; exact-head CI | Evaluation Runtime |
| Continuous Eval Studio | Beta | /online-evals and /online-evals/{monitorId} for draft, simulation, activation, operation, and remediation | Monitor/version CRUD, simulate/activate, pause/resume, bounded backfill/cancel/resume, sample/run/health/topic/candidate reads, alert delivery, and verified remediation operations under /api/online-evals/* | No dedicated TypeScript, Python, or CLI monitor workflow; use authenticated REST | Unit, API, PostgreSQL lifecycle, DOM, and golden-path proof source; exact-head CI | Evaluation Runtime |
| Review Queue | Beta | /review | get_evaluations_id_human_review, post_evaluations_id_human_review, post_evaluations_id_human_review_reviews | CLI review supports local/gate workflows; no complete queue client | Review route tests; exact-head CI | Trust & Review |
| Synthetic and golden lifecycle | Beta | /candidates and evaluation Synthesize/Test Cases panels | get_candidates, get_candidates_id, patch_candidates_id, post_candidates_id_replay, post_candidates_id_promote, and evaluation-scoped review/promotion operations | CLI synthesize, promote, and replay; TypeScript and Python canonical golden-case support | Promotion DB tests; exact-head CI | Trust & Review |
| Model Gateway | Beta | /settings → Model Gateway | get/post_model_gateway_configs, config detail/health/rotation, routing-profile, model-sync, call-ledger, and test-call operations | No dedicated generated gateway client or CLI | Gateway ledger DB tests; exact-head CI | Gateway & Providers |
| Evals-as-Code | Beta | CLI/CI first; evaluation pages show applied results | post_evals_as_code_validate, post_evals_as_code_plan, post_evals_as_code_diff, post_evals_as_code_apply, post_evals_as_code_gate, manifest/apply/drift operations | TypeScript CLI is the primary client; Python parity remains incomplete | Rollback and drift DB tests; exact-head CI | Developer Experience |
| Provider onboarding | Beta | /settings → Provider Keys and Model Gateway | Provider-key REST plus Model Gateway config, health-check, model-sync, and test-call operations | Generated contract types only; no dedicated onboarding convenience client or CLI | Onboarding DOM tests; exact-head CI | Gateway & Providers |
| Calibration control plane | Beta | /calibration, plus evaluation-detail and /llm-judge judge workflows | Anchor-set/version/review, mapping/review/diff/drift, workspace, release-gate, scorer-binding, and existing evaluation calibration-proposal operation families | Judge execution exists in both SDKs; no complete SDK or CLI control-plane workflow | Unit, API, PostgreSQL, DOM, and golden-path proof source; exact-head CI | Trust & Review |
| Evaluator measurement integrity | Experimental | /evaluator-validation measurement workspace, plus run summaries through existing evaluation APIs | Release/run lifecycle, coverage-aware policy sampling, automatic post-run validation, delayed outcome attribution/recomputation, canonical AB/BA judge execution, interval-bearing causal metrics, observation execution, metric reads, and measurement-bound release/Evals-as-Code gates | TypeScript and Python SDK measurement clients, including delayed outcomes; evalgate measurement lists releases/runs, inspects metrics, and executes validation runs | Deterministic unit, DOM, route-inventory, OpenAPI, migration, CLI/SDK, and native PostgreSQL cross-tenant proof in the source release; production browser proof remains release-gated | Evaluation Quality |
| Golden Release / AI regression bot | Experimental | /golden-release candidate inbox, decision header, and publication panel (flag-gated) | Sealed OTEL revisions, canonical measurement revisions, GitHub App webhook/publication/check APIs behind EVALGATE_* flags | No dedicated generated SDK convenience client yet; composite Action and sticky comment remain supported | Unit suites for measurement/github-app/otel/trajectory; DOM golden-release tests; migrations 0106-0108; sandbox GitHub App and production browser proof outstanding | Evaluation Quality |
| Reports | Beta | Evaluation Export/Report actions and /reports | get_reports, post_reports, public verification, and revocation operations | CLI emits regression and machine-readable reports; signed report administration has no complete convenience client | Signing keyring DB tests; exact-head CI | Evidence & Reporting |
| Deployment | Beta | CI and release workflow; no single deployment dashboard | get_releases, post_releases, plus Evals-as-Code gate/apply operations | TypeScript CLI and CI integrations are primary; Python parity remains incomplete | Golden path; exact-head CI | Developer Experience |