# AI PM Competency and Scoring Rubric

## Scoring scale

| Score | Evidence |
|---|---|
| 0 | Cannot explain or gives unsafe/incorrect guidance |
| 1 | Recites terms without applying them |
| 2 | Uses a basic framework but misses evidence, tradeoffs, or closure |
| 3 | Makes a complete, testable decision across user, system, metric, risk, and delivery |
| 4 | Handles counterexamples, uncertainty, organizational effects, and system improvement |

Score only demonstrated answers or artifacts. Record evidence and confidence.

## Ten competency domains

### Strategy and business

Assess market/ICP selection, value, moat, pricing, unit economics, portfolio, and Build/Buy. Evidence: strategy memo, business case, pricing experiment, stop decision.

### User and product discovery

Assess workflow observation, JTBD, baseline, opportunity validation, scope, requirements, and adoption. Evidence: interviews, synthesis, brief, PRD, prototype test.

### AI technology

Assess rules/ML/LLM distinction, Prompt, RAG, tools, Agents, fine-tuning, model selection, failure diagnosis, and architecture tradeoffs. Evidence: architecture proposal and ADR.

### Data

Assess sources, authority, quality, labels, bias, leakage, access, retention, deletion, feedback, and lineage. Evidence: data requirement, schema, sampling and governance plan.

### Evaluation

Assess task taxonomy, golden set, rubric, human/model judging, statistics, error analysis, regression, and online experiment. Evidence: evaluation plan and result.

### AI UX

Assess disclosure, evidence, uncertainty, refusal, clarification, feedback, confirmation, undo, escalation, autonomy, and trust. Evidence: state flow or prototype.

### Project delivery

Assess charter, WBS, phase gates, RACI, RAID, estimates, dependencies, change, vendor, acceptance, handoff, and retrospective. Evidence: project pack and recovery decision.

### Engineering and operations

Assess APIs, observability, versions, latency, capacity, cost, SLO, rollout, rollback, incident, and drift. Evidence: system diagram, runbook, launch review.

### Safety, compliance, and ethics

Assess risk tier, privacy, security, injection, tool authorization, transparency, fairness, copyright, audit, and professional review. Evidence: threat/risk assessment and controls.

### Leadership

Assess clear decisions, writing, conflict, stakeholder mapping, influence, prioritization, escalation, systems, coaching, and portfolio management. Evidence: decision logs, reviews, team mechanisms.

## Graduation levels

- Foundation: all domains at least 1; core terminology and lifecycle understood.
- Contributor: product, technology, evaluation, and delivery at least 2; can complete templates with guidance.
- Independent owner: all domains at least 2; core role domains at least 3; no safety/rollback score below 2.
- Product-line leader: strategy, economics, platform/organization, governance, and leadership at least 3 with real multi-team evidence.

Reading completion is not graduation evidence.

## L0–L5 system mastery mapping

The 0–4 domain scores above evaluate answer and artifact quality. Translate the whole evidence portfolio into the product's mastery ladder as follows:

| Level | Meaning | Minimum evidence |
|---|---|---|
| L0 Entry | Understand the two roles and the lifecycle | Can locate the right workflow and identify obvious unsafe guidance |
| L1 Know | Explain terms, purposes, boundaries, and common failure | All relevant domains at least 1 |
| L2 Do | Produce acceptable artifacts with templates and review | Core role domains at least 2; no safety or rollback domain below 2 |
| L3 Decide | Compare options, quantify tradeoffs, and revise under counterevidence | Core role domains at least 3 with an integrated case |
| L4 Own | Lead a real cross-functional outcome through launch, operations, and failure | Real project result, risk decision, operational evidence, and retrospective |
| L5 Build systems | Create mechanisms that improve decisions across teams without constant personal supervision | Multi-team standard, platform, governance, portfolio, or talent-system evidence |

Do not award L4 or L5 from synthetic cases, reading completion, interview fluency, or self-reported confidence alone. State “readiness estimate” until independent project evidence and review exist.

## Artifact rubric

Evaluate each artifact on:

1. User task and measurable baseline.
2. Evidence quality and labeled assumptions.
3. Alternatives and explicit tradeoffs.
4. Scope, ownership, dependencies, and decisions.
5. Quality metrics and acceptance.
6. Failure, safety, human control, and recovery.
7. Cost, operations, and maintainability.
8. Clarity, traceability, and next action.

Score 0–4 per dimension and require revision for any critical dimension below 2.
