行业知识 · v2.3.0 · 资料核对 2026-10-03
Metrics and AI UX Reference
Metric tree
Use one task-centered outcome plus diagnostic and guardrail metrics.
Business outcome
└─ Qualified task success
├─ Adoption: eligible users who attempt the task
├─ Completion: tasks reaching an outcome
├─ Quality: correct, supported, complete, safe
├─ Efficiency: time, steps, human effort
├─ Reliability: latency, errors, availability
└─ Economics: cost per successful task, margin
Guardrails: severe error, privacy/security event, unwanted action,
complaint, unfair impact, excessive human rework, budget breach
Do not use calls, messages, generated words, or time spent as a North Star unless they directly represent value.
Task success definition
Specify:
- Unit of analysis: request, session, workflow, case, or business outcome.
- Eligible population and exclusions.
- Full, partial, failed, refused, and unsafe outcomes.
- Judge: user, expert, rule, system event, or calibrated model judge.
- Time window and attribution.
- Segment breakdown and confidence/uncertainty.
Baseline
Measure the current manual, rule, search, or old-system process with the same task definition. Include time, cost, error, abandonment, escalation, and user effort. A model benchmark is not a product baseline.
AI UX states
Design all relevant states:
- Ready: explain capability, limits, required data, and examples.
- Working: show progress and allow cancellation for long tasks.
- Result: distinguish source facts, inference, and generated suggestions.
- Low confidence: ask a targeted question, narrow scope, or offer alternatives.
- Refusal: explain the boundary and safe next step.
- Tool action: preview target, parameters, side effects, and reversibility.
- Partial failure: preserve successful work and identify unfinished steps.
- Escalation: transfer context to a human without forcing repetition.
- Correction: let users edit, report, retry, undo, or replace the output.
Autonomy ladder
| Level | Behavior | Control |
|---|---|---|
| L0 | Explain or recommend | User acts |
| L1 | Draft | User edits and submits |
| L2 | Execute after confirmation | Preview and approve |
| L3 | Execute within limits | Budget, scope, monitoring, undo |
| L4 | Plan and act across steps | Strong sandbox, recovery, human supervision |
Reduce autonomy when impact is high, reversibility is low, evidence is weak, or user intent is ambiguous.
Trust design
Build justified trust, not maximum trust. Provide evidence, provenance, freshness, limitations, material uncertainty, action preview, status, and recourse. Avoid decorative confidence scores that are not calibrated.
Feedback design
Collect feedback tied to a task and error type. Combine explicit ratings with edits, retries, abandonment, escalation, and downstream outcomes. Do not reuse feedback for training beyond disclosed purpose and authority.