典 · AI PM 永乐大典

行业知识 · v2.3.0 · 资料核对 2026-10-03

Evaluation Playbook

Decision-first design

Name the decision and candidate versions before choosing metrics. Common decisions: feasibility, model selection, Prompt/index change, vendor acceptance, release, autonomy expansion, or regression response.

Dataset construction

Build strata from production or realistic workflows:

Track source, collection date, authorization, task segment, difficulty, expected behavior, sensitive-data handling, and inclusion rationale. Separate development, validation, and final test sets.

Rubric design

Use observable dimensions. Define each score with positive, negative, and boundary examples. Typical dimensions:

Record severity separately from quality score. Define critical failure conditions as binary gates.

Evaluators

Use deterministic checks for schemas, exact fields, calculations, permissions, tool arguments, tests, latency, and cost. Use domain experts for high-stakes correctness. Use calibrated model judges for scalable comparative review.

Calibrate model judges by:

Metrics by system type

SystemPrimary metrics
Classificationprecision, recall, F1, calibration, error cost
RetrievalRecall@K, MRR, nDCG, permission correctness
RAG answergrounded correctness, citation accuracy, completeness, refusal
Extractionfield-level precision/recall, schema and semantic validity
Codecompile/test/task pass, security, accepted change
Agentend-to-end success, correct tools/args, steps, loops, unsafe action
Recommendationranking, diversity, long-term outcome, fairness

Always add latency, reliability, cost, stability, and severe-tail metrics.

Comparison

Use the same dataset, configuration, limits, and evaluation procedure. Prefer paired comparison on the same tasks. Report sample size, uncertainty, segment results, missing data, judge agreement, and practical effect—not only a single average.

Error analysis

For each important failure, record task, expected/actual, severity, layer, reproducibility, root-cause hypothesis, fix, owner, and regression inclusion. Cluster before prioritizing; prioritize by expected harm and product impact, not only count.

Regression system

Maintain:

Trigger re-evaluation when model, Prompt, data, index, tool, policy, code, user mix, or operating environment changes.