<!-- The AI Transformation Field Guide, v2026.09.1, section: The Enterprise AI Capability Model, https://erikcaldwell.com/field-guide/enterprise-ai-capability-model/ -->

Use this as a scoring instrument, not a reading list. Score your organization honestly against each of the 15 capability areas below. Most organizations at this scale sit at Level 1 or Level 2 in most areas in September 2026, and that is a normal starting point, not a failing grade. The point of the model is to see which areas are furthest behind relative to where your strategy actually needs them, so investment goes where it matters rather than where it's easiest to show progress.

**Level 1: Experimental:** Ad hoc, individual-initiative, no shared standard.
**Level 2: Managed:** Deliberate for the highest-visibility cases, inconsistent elsewhere.
**Level 3: Scaled:** Consistent standard practice across the organization.
**Level 4: AI-Native:** The capability is fully embedded and largely automated; it is simply how the organization operates.

---

## 1. Strategy & Operating Model

- **L1 Experimental:** AI initiatives are ad hoc pilots run by individual teams with no central visibility or prioritization; there is no owned portfolio.
- **L2 Managed:** A central function tracks all AI initiatives, applies a build/buy/configure framework, and pilots have defined owners and kill criteria.
- **L3 Scaled:** AI investment is planned as part of the regular budget cycle under a deliberate hub-and-spoke operating model; multi-model vendor strategy is intentional, not accidental.
- **L4 AI-Native:** AI capability decisions are indistinguishable from core technology strategy; build/buy/configure calls are made against live cost and capability data, not annual planning cycles.

## 2. AI Platform & Gateway Architecture

- **L1:** Teams call model APIs directly and independently; no shared gateway, no consistent access control.
- **L2:** A gateway exists for at least the highest-risk use cases, enforcing basic authentication and rate limits.
- **L3:** All production AI traffic flows through a common gateway enforcing policy, cost limits, and data-loss prevention consistently.
- **L4:** The platform performs dynamic model routing, fallback, and cost-aware model selection automatically, with governed self-service onboarding for new use cases.

## 3. Data & Knowledge Architecture

- **L1:** AI features query raw data sources directly with no retrieval architecture or designed-in permissions enforcement.
- **L2:** RAG is deployed for flagship use cases; permissions are enforced but inconsistently across systems.
- **L3:** Permissions-aware retrieval, data classification, and source provenance are standard requirements for any new AI feature touching enterprise data.
- **L4:** A unified, governed knowledge layer serves every AI system consistently, with automatic staleness checks and access-control validation built in.

## 4. Software Engineering

- **L1:** Developers use AI autocomplete or chat individually; no organizational measurement of impact.
- **L2:** Coding assistants are standard-issue; some teams pilot coding agents for defined tasks under full human review.
- **L3:** Coding agents operate under defined supervision ratios; CI/CD quality gates and architecture validation have been rebuilt for AI-generated code volume.
- **L4:** Agent-directed engineering is standard for well-specified work; supervised software factories run fleets of agents against automated quality gates, and engineers have re-skilled toward specification, review, and system design.

## 5. Workforce Productivity (Knowledge Workers)

- **L1:** Employees use general-purpose chat assistants informally; no enterprise assistant or search is deployed.
- **L2:** An enterprise AI assistant and enterprise search are deployed organization-wide with basic adoption tracking.
- **L3:** Function-specific digital workers (HR, finance, legal, sales, procurement) are deployed for defined workflows with measured task completion.
- **L4:** Multi-step business processes run with agentic assistance end-to-end; humans supervise outcomes rather than perform the underlying tasks.

## 6. Agents & Automation

- **L1:** "Agent" is used loosely for any AI feature; no shared architecture or autonomy-level framework exists.
- **L2:** Agents are deployed for narrow, well-bounded tasks, with human approval required for any consequential action.
- **L3:** An explicit autonomy-level framework governs every production agent; sandboxing, rollback, and idempotency are standard requirements, not afterthoughts.
- **L4:** Agents coordinate across multiple workflows with autonomy calibrated per task type; the organization can state with confidence where agents outperform deterministic software for its own use cases and where they don't.

## 7. Security

- **L1:** AI systems are covered only by general application security practice; no LLM-specific threat model exists.
- **L2:** The OWASP Top 10 for LLM Applications has been reviewed; prompt-injection defenses exist for the highest-exposure systems.
- **L3:** Zero-trust principles are applied specifically to agents; MCP and tool/plugin supply-chain risk is actively managed; AI red teaming happens on a regular cadence.
- **L4:** Continuous adversarial testing runs against production AI systems; the security posture assumes any agent may be compromised and is designed to contain the blast radius by default.

## 8. Identity (Agent & Non-Human Identity)

- **L1:** Agents share service-account credentials with humans or with each other; no distinct agent identity exists.
- **L2:** High-risk agents have distinct credentials, but delegation and scoping are managed manually, case by case.
- **L3:** A non-human identity framework issues every agent its own identity with least-privilege scoping and transaction limits, enforced through delegated, JIT authorization.
- **L4:** Agent identity, authorization, and audit trails are fully standardized platform-wide; any agent's authority can be traced, scoped, and revoked in real time.

## 9. Governance, Legal & Compliance

- **L1:** No formal AI risk framework or system inventory exists; compliance is handled reactively, case by case.
- **L2:** An AI system inventory exists for the highest-visibility systems, with informal risk classification.
- **L3:** NIST AI RMF or ISO/IEC 42001 principles are formally adopted; every production AI system is inventoried and risk-classified; EU AI Act exposure has been assessed regardless of EU footprint.
- **L4:** AI governance is embedded in standard technology governance processes, with automated system registration, risk classification, and compliance monitoring across the full portfolio.

## 10. Evaluation & Quality Engineering

- **L1:** AI feature quality is assessed informally, by demo or anecdote, with no systematic testing.
- **L2:** Golden datasets and basic evals exist for the highest-profile AI features only.
- **L3:** Evals run automatically before every production deployment; hallucination/groundedness checks and trajectory evaluation are standard for agentic systems; red teaming happens on a defined schedule.
- **L4:** Continuous production evaluation runs alongside live traffic, catching quality regressions (including silent model updates from vendors) before they reach end users at scale.

## 11. Observability & AgentOps

- **L1:** AI system behavior is opaque; failures are diagnosed manually and inconsistently.
- **L2:** Logging exists for the highest-risk agents, but there is no standardized tracing across systems.
- **L3:** LLMOps/AgentOps tooling traces agent trajectories consistently; incident-response playbooks exist for AI-specific failures, with audit replay capability.
- **L4:** Full-fidelity tracing, alerting, and audit replay are standard across every production agent, integrated with the organization's broader observability stack.

## 12. AI FinOps & Economics

- **L1:** AI spend is tracked only at the vendor-invoice level, with no per-team or per-use-case visibility.
- **L2:** Cost is tracked per major initiative, but chargeback or showback is informal or absent.
- **L3:** Token economics are understood in cost-per-outcome terms, not just cost-per-token; routing, caching, and model substitution are used deliberately; chargeback/showback is standard practice.
- **L4:** Real-time cost governance catches runaway agent spend automatically; cost-per-outcome is reviewed alongside every other business unit's financials as a matter of course.

## 13. Change Management & Adoption

- **L1:** AI tools are introduced with no formal change management; adoption is left to individual initiative.
- **L2:** Basic role-based AI literacy training exists; an acceptable use policy has been published.
- **L3:** AI champions and communities of practice operate across business units; a use-case library captures and spreads what's working; displacement fears are addressed directly rather than avoided.
- **L4:** AI fluency is a standard part of role competency expectations organization-wide; adoption is measured and managed like any other major change program, with continuous feedback loops.

## 14. AI-Native Product Development

- **L1:** AI features are bolted onto existing chatbot-style interfaces with no distinct design discipline.
- **L2:** Some products incorporate citations or confidence indicators for AI-generated content, inconsistently.
- **L3:** Progressive autonomy and approval UX are designed deliberately; failure states, provenance, and confidence indicators are first-class design considerations across products.
- **L4:** Products are designed AI-native from the outset (including, where relevant, treating AI agents as users of the product's own interfaces) with autonomy and trust calibrated by design rather than added afterward.

## 15. Measurement & Business Value

- **L1:** AI success is reported through activity metrics: seats provisioned, prompts sent, "AI-assisted" volume.
- **L2:** Some engineering metrics (DORA-style) are tracked for AI-assisted development, but knowledge-worker and enterprise value metrics remain activity-based.
- **L3:** Outcome metrics (revenue enabled, cost reduced, capacity released) are tracked for major initiatives, and skill-segmented productivity effects are understood rather than assumed uniform.
- **L4:** AI value measurement is integrated into standard business reporting; every material AI investment has an owned outcome metric the organization would defend to its board without relying on adoption or usage numbers.

---

**Using this model:** Score each area today, then again every two quarters: maturity here moves in months, not years, in either direction. Prioritize investment in whichever areas are most behind relative to your strategy's actual dependencies, not the areas easiest to show a Level 4 slide about. A common and costly failure pattern is racing to Level 3-4 in Workforce Productivity or AI-Native Product Development while Security, Identity, and Governance sit at Level 1: visible progress with an invisible, compounding liability underneath it.
