行业知识 · v2.3.0 · 资料核对 2026-10-03
Evaluation Playbook
Decision-first design
Name the decision and candidate versions before choosing metrics. Common decisions: feasibility, model selection, Prompt/index change, vendor acceptance, release, autonomy expansion, or regression response.
Dataset construction
Build strata from production or realistic workflows:
- Core/high-frequency tasks.
- High-value tasks.
- Known failures and regression cases.
- Long-tail languages, formats, lengths, and user groups.
- No-answer, ambiguous, conflicting, and outdated-context cases.
- Adversarial, misuse, injection, permission, and privacy cases.
- High-impact cases where error severity dominates frequency.
Track source, collection date, authorization, task segment, difficulty, expected behavior, sensitive-data handling, and inclusion rationale. Separate development, validation, and final test sets.
Rubric design
Use observable dimensions. Define each score with positive, negative, and boundary examples. Typical dimensions:
- Task success and instruction adherence.
- Factual correctness and evidence support.
- Completeness and relevance.
- Correct refusal and escalation.
- Permission and tool correctness.
- Safety and policy compliance.
- Style only when it affects task value.
Record severity separately from quality score. Define critical failure conditions as binary gates.
Evaluators
Use deterministic checks for schemas, exact fields, calculations, permissions, tool arguments, tests, latency, and cost. Use domain experts for high-stakes correctness. Use calibrated model judges for scalable comparative review.
Calibrate model judges by:
- Comparing with a human-labeled gold subset.
- Randomizing answer order and hiding model identity.
- Testing position, length, style, and self-preference bias.
- Reviewing human/judge disagreements.
- Recalibrating when task, rubric, or judge changes.
Metrics by system type
| System | Primary metrics |
|---|---|
| Classification | precision, recall, F1, calibration, error cost |
| Retrieval | Recall@K, MRR, nDCG, permission correctness |
| RAG answer | grounded correctness, citation accuracy, completeness, refusal |
| Extraction | field-level precision/recall, schema and semantic validity |
| Code | compile/test/task pass, security, accepted change |
| Agent | end-to-end success, correct tools/args, steps, loops, unsafe action |
| Recommendation | ranking, diversity, long-term outcome, fairness |
Always add latency, reliability, cost, stability, and severe-tail metrics.
Comparison
Use the same dataset, configuration, limits, and evaluation procedure. Prefer paired comparison on the same tasks. Report sample size, uncertainty, segment results, missing data, judge agreement, and practical effect—not only a single average.
Error analysis
For each important failure, record task, expected/actual, severity, layer, reproducibility, root-cause hypothesis, fix, owner, and regression inclusion. Cluster before prioritizing; prioritize by expected harm and product impact, not only count.
Regression system
Maintain:
- Small blocking core set for every change.
- High-risk gate set.
- Capability-specific suites.
- Historical fixed-failure set.
- Periodically refreshed production sample.
Trigger re-evaluation when model, Prompt, data, index, tool, policy, code, user mix, or operating environment changes.