行业知识 · v2.3.0 · 资料核对 2026-10-03
AI System Failure Diagnostics
Triage order
- Confirm the expected behavior and reproduce with pinned versions.
- Determine scope by task, user, tenant, model, version, region, and time.
- Check data/telemetry correctness before interpreting metrics.
- Locate the first layer where observed behavior diverges.
- Apply the smallest safe fix and add a regression case.
Layer map
| Symptom | Check first | Common fixes |
|---|---|---|
| Wrong user intent | input normalization, missing clarification, task routing | clarify, constrain task, improve routing |
| Missing fact | source availability, sync, ACL, chunking, recall | ingest/fix ACL, retrieval, query rewrite |
| Retrieved evidence but wrong answer | context order/noise, instruction, model adherence | rerank, trim, structure, validate, change model |
| Correct text, wrong tool action | schema, arguments, mapping, authorization | typed schema, validation, confirmation |
| Duplicate action | retry/idempotency/state | idempotency key, state machine, reconciliation |
| Agent loop | no-progress detection, tool errors, budget | stop rules, planner constraints, handoff |
| JSON valid but semantically wrong | only syntax checked | domain validation, constraints, human review |
| Cross-user data | ACL, cache key, tenant context | stop service, isolate, purge, investigate |
| High latency | context, retrieval, model, serial tools, queue | reduce context, parallelize, route, cache safely |
| Cost spike | retries, loops, long input/output, traffic mix | budgets, route, compress, cache, cap |
| Quality regression | model/Prompt/index/tool/data version | rollback, bisect, segment, add regression |
RAG decomposition
Ask in order:
- Does the authorized source contain the answer?
- Was the latest source ingested and deletion propagated?
- Did retrieval return relevant content within K?
- Did reranking place it high enough?
- Did context assembly preserve it without truncation?
- Did the model use it, cite it accurately, and avoid unsupported claims?
- Did the interface display source/freshness correctly?
Measure retrieval and answer generation separately.
Prompt change discipline
Before changing a Prompt, record the failing sample, current version, expected behavior, error class, hypothesis, change, targeted test, full regression result, latency/cost effect, and rollout plan. Avoid adding exceptions indefinitely; redesign the contract when instructions conflict.
Severity
- Critical: unauthorized disclosure/action, severe real-world harm, systemic unsafe behavior.
- High: incorrect high-value decision, repeated action, major customer workflow failure.
- Medium: recoverable task failure or substantial rework.
- Low: style, minor completeness, or low-impact inconvenience.
Frequency does not override severity. A rare critical failure can block release.