Instrument
A7 Agentic AI Readiness Assessment
The A7 readiness assessment scores an organization from 1 to 5 on seven dimensions, from data architecture to autonomy calibration, for a total between 7 and 35. The total names the most autonomous class of AI system the organization can safely run, and the floor rule lowers that level whenever one dimension is too weak to hold it.
Last verified
Scoring
| Dimension | The question it answers | Levels 1 to 5 |
|---|---|---|
| A1 Data Architecture | Can your data serve agents in real time, with the business context they need to decide? | Siloed, Consolidated, Integrated, Adaptive, Agent-optimized |
| A2 Technical Infrastructure | Does your infrastructure support multi-agent coordination, reversible actions and graceful failure? | Basic, Standardized, Orchestrated, Resilient, Agent-native |
| A3 Governance Framework | Does your AI governance cover autonomous decisions, beyond model deployment and data privacy? | Absent, Informal, Structured, Comprehensive, Embedded |
| A4 Human Oversight Protocols | Are supervision, approval workflows, escalation paths and kill switches in place for agent decisions? | None, Minimal, Defined, Graduated, Adaptive |
| A5 Organizational Readiness | Do you have the culture, skills, change management and executive alignment to run agents? | Resistant, Aware, Engaged, Committed, Agent-native |
| A6 Security & Safety | Do you have agent-specific security beyond traditional application security? | Standard AppSec, Aware, Defensive, Proactive, Comprehensive |
| A7 Autonomy Calibration | Can you tell genuine agents from rebranded ones and match autonomy to readiness per use case? | Uncalibrated, Aware, Defined, Calibrated, Optimized |
Score each dimension against the full rubric below, then add the seven scores. The lowest possible total is 7 and the highest is 35.
| Total | Band | What it means |
|---|---|---|
| 7-14 | L0 Not Ready | Traditional AI only: predictive models, recommendation engines, dashboards. No system that takes action on its own. |
| 15-21 | L1 Copilot Ready | AI suggests, drafts and surfaces insight. A human approves every action. |
| 22-28 | L2 Supervised Agent Ready | Agents act within defined boundaries while people monitor their decisions and can step in. |
| 29-33 | L3 Autonomous Ready | Agents run within programmatic guardrails and people review the exceptions. Needs mature governance, security and oversight together. |
| 34-35 | L4 Full Autonomy Ready | Self-directed agents with minimal oversight. Rare, and only for narrow, well-bounded use cases. |
Floor rule
L2 needs every dimension at 2 or higher, L3 every dimension at 3 or higher, and L4 every dimension at 4 or higher; L1 has no floor. Your level is the highest one that both your total and every single dimension support, so a total of 28 with Governance (A3) at 1 is L1, not L2.
Download: A7 Agentic AI Readiness Worksheet
The printable worksheet: every dimension, the score bands and room for your team's evidence.
Enter your email to get instant access. You'll also receive the fortnightly newsletter.
Free. No spam. Unsubscribe anytime.
Why this instrument exists, and what the scores mean for a leadership team: The A7 Framework: Are You Ready for Agentic AI?
How to run it
Score it in one working session with the people who own each dimension: engineering, data, security, governance or risk, AI operations, a business unit lead and an executive sponsor. A team from one function misses the gaps the other functions can see.
- Score each dimension from 1 to 5 against the rubric below, and write down the evidence behind each score.
- Add the seven scores and read the band for your total.
- Apply the floor rule. If your weakest dimension cannot hold the band, your level drops to the one it can hold.
- Compare that level with the most autonomous system you already run. Running above your level is premature autonomy, the pattern this instrument exists to catch.
- Run the Agent Washing Detector on every system you call an agent.
- Plan the sprint: fix the floor first, then raise the weakest dimensions.
A worked case: a total of 24 sits in the L2 band. If Human Oversight (A4) scores 1, the floor rule holds the organization at L1, and all improvement effort goes to A4 until it reaches 2.
The full rubric follows, one dimension at a time. Pick the level that matches what runs today, not what is on the roadmap.
A1 Data Architecture
Can your data infrastructure serve agents in real time with the business context they need to make decisions?
- Siloed. Data sits in departmental silos and moves in nightly batches. There is no API access to operational data, training data is pulled by hand, and nobody keeps a catalog or lineage.
- Consolidated. A basic warehouse or lake exists with some API endpoints, but quality is uneven, metadata is thin and each new use case needs its own integration work.
- Integrated. An integrated lakehouse offers broad API access. The team uses a data catalog, tracks the core quality metrics and serves near-real-time data for priority systems, and a semantic layer is taking shape.
- Adaptive. Data streams in real time through a full semantic layer that carries business context. Domain teams own and publish their data products, and agents find what they need through self-service interfaces.
- Agent-optimized. The data estate is built for agents: it streams in real time, exposes agent-specific interfaces, enforces quality automatically and learns from the feedback that agent actions generate.
A2 Technical Infrastructure
Does your compute, networking and orchestration support multi-agent coordination, reversible actions and graceful failure?
- Basic. Workloads run in one cloud environment and are deployed by hand. There is no container orchestration or agent-specific infrastructure, and monitoring stops at the application.
- Standardized. Workloads are containerized with basic CI/CD, but agents run as standalone processes with no way to undo an action and nothing to coordinate several agents.
- Orchestrated. Containers are orchestrated and auto-scale, an agent orchestration framework is in place with basic state management, and the team monitors agent execution.
- Resilient. The platform spans regions or clouds and manages how agents use tools. Critical actions can be reversed, failover is automatic, every run leaves an end-to-end trace and performance is held to SLAs.
- Agent-native. A multi-cloud agent mesh routes work dynamically and every operation is atomic, with automatic rollback. Capacity scales with agent load, agents are versioned and retired like software, and the infrastructure heals itself.
A3 Governance Framework
Does your AI governance cover autonomous decision-making as well as model deployment and data privacy?
- Absent. There is no AI governance structure, policy, inventory or named owner, so agents are deployed by whoever has access.
- Informal. Basic AI policies exist but nobody enforces them. Governance is one person's part-time job, agents are governed like any other model and nobody has defined what an agent may decide.
- Structured. A complete AI and agent inventory, named governance owners, risk tiers, deployment gates, regular reviews and documented agent decision boundaries. This is Minimum Viable Governance.
- Comprehensive. A cross-functional council holds real authority, governance is built into the agent development lifecycle, each use case has explicit decision rights and the organization can support an external audit.
- Embedded. Governance runs as code: the platform enforces agent policies, monitors decisions against them and reports to the board in real time. At this level governance speeds deployment up.
A4 Human Oversight Protocols
Are supervision models, approval workflows, escalation paths and kill switches in place for agent decisions?
- None. Nobody oversees what agents do. There are no approval workflows, escalation paths, kill switches or monitoring of agent actions.
- Minimal. Agent actions are logged but rarely reviewed, escalation is informal and stopping an agent takes an engineer because there is no self-service kill switch.
- Defined. Critical actions (financial, customer-facing, data changes) pass an approval gate. Logs are reviewed on a schedule, escalation paths are written down and any agent can be paused.
- Graduated. Each agent gets the autonomy its risk tier allows. People watch in real time, one click stops an agent without breaking the service, and anomaly detection calls in a person.
- Adaptive. Oversight tightens or relaxes with confidence, novelty and risk signals. Shutdown is instant and graceful, oversight data feeds governance and the team rehearses failures in regular simulations.
A5 Organizational Readiness
Does the organization have the culture, skills, change management and executive alignment to deploy and manage autonomous agents?
- Resistant. No AI culture, open or quiet resistance, no skills development and no executive sponsor. AI is seen as a threat.
- Aware. Executives talk about AI but have not funded a strategy, skills are thin and change management is ad hoc.
- Engaged. Leadership understands AI and has given it a budget. Skills programs are running, and there is a dedicated AI team, a change playbook and a named executive sponsor.
- Committed. The C-suite agrees on an agentic strategy and funds a dedicated agent operations team to run it. AI literacy reaches across the organization and incentives reward adoption.
- Agent-native. Processes are designed with agents as participants. Every function has agent integration expertise. The board owns the agent strategy.
A6 Security & Safety
Does the organization have agent-specific security beyond traditional application security?
- Standard AppSec. Security is traditional application security. Nobody has assessed prompt injection, tool-use abuse or credential exposure, and there are no agent-specific measures.
- Aware. The security team knows the agent risks and has added basic input validation. Agent credentials sit in a manager instead of the code, though nobody has tested agents adversarially yet.
- Defensive. Prompt injection defenses are in place, tool-use boundaries are enforced and each agent has its own credentials. Customer-facing outputs are validated and the setup is reviewed regularly.
- Proactive. The team keeps an agent threat model current, red-teams regularly, runs untested actions in a sandbox and monitors agent behavior. Security is part of the development lifecycle.
- Comprehensive. A full program: scope control, adversarial testing, sandboxing, credential isolation, output validation, behavioral monitoring and supply-chain security for tools and plugins.
A7 Autonomy Calibration
Can the organization assess its own readiness, tell genuine agents from agent-washed products, and run the right autonomy level per use case?
- Uncalibrated. The organization cannot tell assistants, copilots and agents apart, has no autonomy taxonomy and takes vendor claims at face value.
- Aware. Knows autonomy levels exist, but the taxonomy is informal and varies across teams. Some vendor claims are questioned, with no method behind it.
- Defined. One autonomy taxonomy applies across the organization, vendor evaluation includes a technical assessment and every current deployment is classified correctly.
- Calibrated. Autonomy is matched to readiness for each use case. Every deployment starts with a readiness assessment, scores are revisited regularly and the Agent Washing Detector is in routine use.
- Optimized. Each use case runs at the right autonomy and is reassessed continuously. The path to the next level is written down, along with what it would take to get there.
The Agent Washing Detector
Gartner calls the rebranding of AI assistants, RPA and chatbots as agents "agent washing", and estimates that only about 130 of the thousands of agentic AI vendors are real. Ask these five questions of every vendor product and internal system you call an agent, and score A7 with the answers in hand.
- Planning. Does it break a goal into sub-tasks, or follow a fixed script?
- Tool use. Does it choose and call tools as the task needs, or call one predetermined API?
- Memory. Does it keep state across interactions and learn from earlier actions, or start fresh every time?
- Autonomy. Does it act without a human approving each step? If every step needs a click, it is a copilot.
- Recovery. When something unexpected happens, does it change its plan, or fail and escalate?
Three or more answers of no means the system is most likely agent-washed. A well-built copilot still has value; it should be scored and governed as a copilot.
The Agentic Readiness Sprint
A score becomes a plan in three phases. The logic is the same as paying down debt: clear the gap that blocks the next level before improving anything else.
Phase 1: assess and find the floor (weeks 1 to 2). Run the assessment, apply the floor rule and list every dimension below the minimum for your target level. Give each gap an owner, a budget and 30, 60 and 90 day milestones, and re-score those dimensions at 90 days.
Phase 2: fix the floor (weeks 3 to 12). Put all improvement effort into the dimensions that break the floor rule for your target level, and nothing else. Typical first moves: for A3 at 1, draft basic AI policies, name a governance owner and start an AI system inventory; for A4 at 1, log agent actions, set an escalation path and add a way to intervene by hand; for A6 at 1, brief the security team on agent risks, move agent credentials into a secrets manager and add input validation.
Phase 3: level up (months 4 to 12). With the floor clear, raise the lowest scores first. Moving a dimension from 2 to 3 does more for your level than moving another from 4 to 5. Pilot the next level on selected use cases and check the assessment against what happens.
Reassess the dimensions you are improving every quarter, run the full assessment twice a year, and run it again after a major incident, a significant technology change or a new agent deployment.
Sources
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner press release, 2025-06-25
Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →
Published by AI Exponent LLC