行业知识 · v2.3.0 · 资料核对 2026-10-03
AI Evaluation and Launch Plan
在交互知识库中阅读下载 Markdown 原文
1. Decision
- System / version / candidate:
- Decision: feasibility / selection / acceptance / release / expansion / regression:
- Baseline:
- Decision owner and reviewers:
2. Task and risk taxonomy
Include core, high-value, long-tail, refusal, adversarial, permission, and high-impact cases.
3. Dataset governance
- Source, date, authorization, and sensitive-data treatment:
- Development / validation / test split:
- Leakage controls:
- Version and change log:
- Refresh and retirement:
4. Rubric and evaluators
- Human qualification and adjudication:
- Model-judge calibration and agreement:
- Deterministic checks:
5. Metrics and gates
6. Error analysis
7. Security, privacy, and red-team checks
- Data authority and deletion:
- Tenant/access isolation:
- Prompt injection and untrusted content:
- Tool authorization and irreversible action:
- Abuse, unsafe output, fairness, transparency:
- Professional/legal reviews required:
8. Operational readiness
- Trace/version coverage:
- Dashboards and alerts:
- Capacity, rate limits, supplier/fallback:
- Human sampling and coverage:
- Runbook, support, incident roles:
- Rollback mechanism tested:
9. Rollout stages
10. Decision
- Go / conditional Go / hold / rollback / stop:
- Evidence:
- Violated gates:
- Conditions, owners, due dates, and expiration:
- Rollback triggers:
- Next evaluation trigger: