典 · AI PM 永乐大典

行业知识 · v2.3.0 · 资料核对 2026-10-03

Evaluate and Launch AI

Turn “the demo looks good” into reproducible evidence and a controlled operating decision.

Operating rules

Workflow

1. Define the decision

State what decision the evaluation must support: choose a model, validate a feature, accept a vendor, release a version, expand autonomy, or diagnose regression. Record the current baseline and candidate versions.

2. Build the evaluation specification

Read references/evaluation-playbook.md. Define task taxonomy, sampling, dataset splits, rubric, severity, automated checks, human review, judge calibration, uncertainty, and statistical comparison.

3. Cover the full quality surface

Measure task success, correctness, evidence support, completeness, refusal, safety, permission correctness, tool behavior, latency, cost, stability, and user outcome. Add domain-specific criteria rather than forcing every task into one generic score.

4. Analyze errors

Create an error taxonomy, inspect disagreements and severe failures, identify the failing system layer, and add important fixed failures to the regression set. Report both aggregate and segment results.

5. Build the release gate

Read references/safety-launch-operations.md. Define mandatory approvals, security/privacy tests, operational readiness, traffic stages, monitoring, alerts, human coverage, rollback mechanism, and incident ownership.

6. Make the release decision

Choose Go, conditional Go, hold, rollback, or stop. Tie every condition to an owner, deadline, validation, and expiration. Do not average away a violated hard gate.

7. Produce the artifact

Copy assets/evaluation-and-launch-plan.md. Use assets/golden-set-schema.csv as the minimum row schema for a local evaluation set.

Quality gate

Confirm that the plan includes:

If a hard gate cannot be tested, classify the release as unverified rather than passing it by assumption.