# Metrics and AI UX Reference

## Metric tree

Use one task-centered outcome plus diagnostic and guardrail metrics.

```text
Business outcome
└─ Qualified task success
   ├─ Adoption: eligible users who attempt the task
   ├─ Completion: tasks reaching an outcome
   ├─ Quality: correct, supported, complete, safe
   ├─ Efficiency: time, steps, human effort
   ├─ Reliability: latency, errors, availability
   └─ Economics: cost per successful task, margin

Guardrails: severe error, privacy/security event, unwanted action,
complaint, unfair impact, excessive human rework, budget breach
```

Do not use calls, messages, generated words, or time spent as a North Star unless they directly represent value.

## Task success definition

Specify:

- Unit of analysis: request, session, workflow, case, or business outcome.
- Eligible population and exclusions.
- Full, partial, failed, refused, and unsafe outcomes.
- Judge: user, expert, rule, system event, or calibrated model judge.
- Time window and attribution.
- Segment breakdown and confidence/uncertainty.

## Baseline

Measure the current manual, rule, search, or old-system process with the same task definition. Include time, cost, error, abandonment, escalation, and user effort. A model benchmark is not a product baseline.

## AI UX states

Design all relevant states:

- Ready: explain capability, limits, required data, and examples.
- Working: show progress and allow cancellation for long tasks.
- Result: distinguish source facts, inference, and generated suggestions.
- Low confidence: ask a targeted question, narrow scope, or offer alternatives.
- Refusal: explain the boundary and safe next step.
- Tool action: preview target, parameters, side effects, and reversibility.
- Partial failure: preserve successful work and identify unfinished steps.
- Escalation: transfer context to a human without forcing repetition.
- Correction: let users edit, report, retry, undo, or replace the output.

## Autonomy ladder

| Level | Behavior | Control |
|---|---|---|
| L0 | Explain or recommend | User acts |
| L1 | Draft | User edits and submits |
| L2 | Execute after confirmation | Preview and approve |
| L3 | Execute within limits | Budget, scope, monitoring, undo |
| L4 | Plan and act across steps | Strong sandbox, recovery, human supervision |

Reduce autonomy when impact is high, reversibility is low, evidence is weak, or user intent is ambiguous.

## Trust design

Build justified trust, not maximum trust. Provide evidence, provenance, freshness, limitations, material uncertainty, action preview, status, and recourse. Avoid decorative confidence scores that are not calibrated.

## Feedback design

Collect feedback tied to a task and error type. Combine explicit ratings with edits, retries, abandonment, escalation, and downstream outcomes. Do not reuse feedback for training beyond disclosed purpose and authority.
