---
name: design-ai-solution
description: Design testable AI product architectures and make explicit choices among rules, traditional machine learning, Prompt engineering, RAG, fine-tuning, tools, workflows, and Agents. Use for model or vendor selection, technical solution design, architecture reviews, RAG or Agent planning, context engineering, cost and latency tradeoffs, failure diagnosis, and architecture decision records.
---

# Design AI Solution

Design the smallest reliable system that can meet the product contract. Prefer deterministic components for deterministic needs and reserve model autonomy for irreducible uncertainty.

## Operating rules

- Start from task, quality, risk, latency, cost, privacy, and scale constraints.
- Treat model output as one component, not the complete system.
- Do not choose a model from public rankings alone; require task-specific evaluation.
- Do not prescribe Agent architecture when a fixed workflow can solve the problem.
- Enforce authorization outside the model. Never rely on Prompt text as the only security boundary.
- Separate architecture facts, assumptions, tradeoffs, and experiments.
- Keep external calls and implementation changes out of scope unless the user explicitly requests them.

## Workflow

### 1. Read the product contract

Extract task types, inputs, outputs, knowledge freshness, tool actions, permissions, risk, autonomy, quality thresholds, traffic, latency, cost, and deployment constraints. If these are missing, state the assumptions needed for a provisional design.

### 2. Partition the problem

Split the system into deterministic workflow, traditional ML, generative model, retrieval, tools, human decision, and observability. Assign AI only to steps where uncertainty or unstructured information justifies it.

### 3. Choose the capability pattern

Read [references/architecture-decisions.md](references/architecture-decisions.md) and decide among Prompt-only, RAG, tool calling, fixed workflow, bounded Agent, fine-tuning, traditional ML, or a hybrid. Explain rejected options.

### 4. Define the end-to-end design

Specify:

- Input normalization and trust boundaries.
- Context construction, memory, retrieval, and freshness.
- Model candidates, routing, fallback, and version pinning.
- Tool schemas, credentials, authorization, idempotency, and confirmation.
- Output schema, validation, citations, refusal, and human handoff.
- Logging, tracing, privacy, metrics, budgets, and recovery.

### 5. Build a test plan before implementation

Map each critical design decision to an experiment, metric, representative data, threshold, and decision rule. Include quality, high-risk tails, latency, cost, capacity, and operational failure.

### 6. Diagnose existing failures

Read [references/failure-diagnostics.md](references/failure-diagnostics.md). Locate the failing layer before changing Prompt or model. Separate data, retrieval, model, tool, orchestration, interface, permission, and observability causes.

### 7. Produce the artifact

Copy [assets/solution-design-template.md](assets/solution-design-template.md) for a solution proposal. Use [assets/architecture-decision-record.md](assets/architecture-decision-record.md) for a discrete architecture choice.

## Quality gate

Confirm that the design states:

- The non-AI baseline and simpler alternative.
- Every component's responsibility and failure behavior.
- The source, authorization, freshness, and deletion path for data/context.
- Why RAG, tools, fine-tuning, or Agent autonomy is necessary.
- Model-independent interfaces and a practical fallback.
- Quality, latency, throughput, availability, and per-successful-task cost targets.
- High-risk action controls, human confirmation, audit, and undo/recovery.
- Versioning for model, Prompt, index, tools, code, and evaluation set.
- Experiments that can falsify the design.

Reject “use a stronger model” as a complete architecture recommendation.
