# Architecture Decision Reference

## Selection order

1. Remove unnecessary steps or solve with process design.
2. Use deterministic code/rules for exact logic.
3. Use search or database query for direct retrieval.
4. Use traditional ML for stable classification, ranking, forecasting, or anomaly detection.
5. Use an LLM for semantic interpretation, synthesis, generation, or flexible extraction.
6. Add RAG, tools, workflow, Agent autonomy, or fine-tuning only for a diagnosed need.

## Pattern chooser

| Need | Pattern | Main evaluation |
|---|---|---|
| Stable behavior, format, examples | Prompt + constrained output | adherence, semantic correctness |
| Dynamic/private/evidenced knowledge | RAG | retrieval, grounded answer, permission |
| Live information or side effects | Tool calling | correct tool/args, authorization, idempotency |
| Known multi-step process | Deterministic workflow | step success, recovery, observability |
| Dynamic path based on intermediate results | Bounded Agent | task success, steps, loops, unsafe actions |
| Repeated behavior not solved by Prompt | Fine-tuning | held-out gain, regression, maintenance cost |
| Numeric prediction/ranking at scale | Traditional ML | calibration, ranking/prediction metrics, drift |

Combine patterns only when each component has a clear responsibility.

## RAG design

Specify ingestion, parsing, deduplication, chunking, metadata, ACL, index, query rewrite, retrieval, reranking, context assembly, generation, citation, no-answer, sync, deletion, and evaluation.

Key tradeoffs:

- Smaller chunks improve precision but may lose meaning.
- More retrieved context improves recall but can add noise, latency, and cost.
- Hybrid lexical/vector retrieval often handles names and concepts better than either alone.
- Access control must occur before content reaches the model.
- Citation presence is not citation correctness.

## Tool and Agent design

Keep credentials outside the model. Validate tool name, arguments, user/tenant authorization, rate, budget, target, and side effects in deterministic code. Use preview/confirmation for high-impact or irreversible actions.

Set limits for steps, time, tokens, money, retries, tool errors, duplicate state, and lack of progress. Preserve intermediate state for recovery and audit.

## Fine-tuning decision

Fine-tune only after establishing a Prompt/RAG baseline and a stable evaluation set. Confirm that the problem is repeated behavior or capability adaptation, not missing current facts, broken retrieval, ambiguous requirements, or poor tool design.

Track data origin, rights, quality, train/test leakage, model version, hyperparameters, deployment, rollback, and regression.

## Model and vendor selection

Use hard filters first: input/output modality, language, context, tools, deployment, data policy, region, safety, SLA, and compatibility. Then compare on owned tasks:

- Core and high-risk quality.
- Stability and refusal.
- P50/P95 latency and throughput.
- Total cost per successful task.
- Rate limits, availability, support, and change policy.
- Portability and fallback.

## Context engineering

Allocate context deliberately among system policy, task instruction, user data, retrieved knowledge, tool state, memory, and output budget. Remove repeated/noisy content, summarize with provenance, and protect instruction hierarchy from untrusted text.

Store memory only when authorized and useful beyond the current turn. Attach source, scope, timestamp, confidence, retention, edit, and deletion controls.

## Observability and versioning

Trace request, user/tenant, feature, input class, model, Prompt, retrieval/index, tool calls, output validator, latency, tokens/cost, safety decisions, feedback, and final outcome. Redact or minimize sensitive data.

Version model, Prompt, configuration, retriever, index, tools, code, policies, and evaluation set so every production result can be reconstructed.
