---
name: evaluate-launch-ai
description: Build AI evaluation systems and controlled release plans covering golden datasets, rubrics, offline and online metrics, human and model judging, red teaming, safety gates, launch reviews, canaries, monitoring, rollback, incident response, and regression control. Use for evaluation design, model comparison, AI acceptance criteria, Go or No-Go reviews, production launch, quality regressions, and post-launch governance.
---

# Evaluate and Launch AI

Turn “the demo looks good” into reproducible evidence and a controlled operating decision.

## Operating rules

- Evaluate the end-to-end task, not only model text.
- Sample real tasks and deliberately include long-tail, refusal, adversarial, and high-impact cases.
- Keep the test set versioned and protected from tuning leakage.
- Separate average quality from severe tail risk.
- Calibrate LLM judges against human gold labels; do not treat them as unquestionable ground truth.
- Define launch, rollback, and stop thresholds before viewing final results.
- Never weaken a safety or privacy gate to make a release appear ready.

## Workflow

### 1. Define the decision

State what decision the evaluation must support: choose a model, validate a feature, accept a vendor, release a version, expand autonomy, or diagnose regression. Record the current baseline and candidate versions.

### 2. Build the evaluation specification

Read [references/evaluation-playbook.md](references/evaluation-playbook.md). Define task taxonomy, sampling, dataset splits, rubric, severity, automated checks, human review, judge calibration, uncertainty, and statistical comparison.

### 3. Cover the full quality surface

Measure task success, correctness, evidence support, completeness, refusal, safety, permission correctness, tool behavior, latency, cost, stability, and user outcome. Add domain-specific criteria rather than forcing every task into one generic score.

### 4. Analyze errors

Create an error taxonomy, inspect disagreements and severe failures, identify the failing system layer, and add important fixed failures to the regression set. Report both aggregate and segment results.

### 5. Build the release gate

Read [references/safety-launch-operations.md](references/safety-launch-operations.md). Define mandatory approvals, security/privacy tests, operational readiness, traffic stages, monitoring, alerts, human coverage, rollback mechanism, and incident ownership.

### 6. Make the release decision

Choose Go, conditional Go, hold, rollback, or stop. Tie every condition to an owner, deadline, validation, and expiration. Do not average away a violated hard gate.

### 7. Produce the artifact

Copy [assets/evaluation-and-launch-plan.md](assets/evaluation-and-launch-plan.md). Use [assets/golden-set-schema.csv](assets/golden-set-schema.csv) as the minimum row schema for a local evaluation set.

## Quality gate

Confirm that the plan includes:

- A representative task taxonomy and documented sampling gaps.
- Clear definitions of success, partial success, failure, and critical failure.
- Rubrics with examples and an adjudication process.
- A baseline, candidate comparison, uncertainty, and segment analysis.
- Retrieval/tool/system checks where applicable.
- Security, privacy, misuse, fairness, and human-control tests proportional to risk.
- Explicit quality, safety, latency, cost, and reliability thresholds.
- Canary stages, observability, rollback, incident response, and post-launch sampling.
- Regression ownership and triggers for re-evaluation.

If a hard gate cannot be tested, classify the release as unverified rather than passing it by assumption.
