Agentic AI becomes useful when it removes coordination work without removing accountability. That distinction matters because many organisations are currently evaluating agents through the wrong lens. They ask whether a model can perform a task. The more important question is whether the operating process can safely absorb machine initiative.
An impressive demonstration can hide a weak operating design. A model may summarise a case, classify a document, search a knowledge base and even invoke an API. None of those capabilities tell you whether it should be allowed to decide eligibility, move money, close a regulatory case or change a production system. Capability and authority are different design decisions.
Start with the work that surrounds the decision
In large operating environments, the expensive part of a decision is often not the decision itself. It is the preparation. Someone has to notice that work has arrived, assemble the relevant history, check whether information is missing, compare the case against policy, identify duplicates, find the right owner, prepare a recommendation and record the evidence. The final judgement may take five minutes after forty minutes of coordination.
That is where agentic patterns become interesting. The agent is not valuable because it imitates a person. It is valuable because it can continuously perform bounded coordination across systems and queues. In an emergency-assistance process, for example, an AI-enabled workflow might assemble the application, confirm that identity checks have been completed, compare the address with an impacted-area polygon, identify related household applications, retrieve the applicable programme rules and prepare a case brief. A human assessor can then make the consequential decision with much better context.
This is a materially different proposition from allowing a model to autonomously approve or reject assistance. One design compresses administrative latency. The other transfers decision authority.
A five-part test for agent suitability
I use five questions before talking about models or products.
- Signal: Is there a reliable event that starts the work? A new application, failed test, queue threshold, changed record or operational alert is a better trigger than an ambiguous instruction to “watch for problems”.
- Context: Can the agent retrieve authoritative information rather than merely plausible information? Context should come from governed systems of record, approved knowledge and explicit tools.
- Policy: Are the constraints legible enough to encode or test? If experienced staff cannot agree on the rule, automating it will not resolve the ambiguity.
- Action: Is the proposed action reversible, observable and proportionate to the confidence available? Drafting is different from sending. Recommending is different from approving.
- Evidence: Can the system show what information, rule and tool invocation produced the outcome? In consequential workflows, “the model said so” is not evidence.
This test deliberately shifts the conversation away from model capability and toward operating control. It also exposes the cases where AI should remain assistive. If context is incomplete or the policy is inherently discretionary, the right design may be to improve information preparation rather than automate judgement.
The autonomy ladder is more useful than the automation percentage
Executives are often shown a percentage of work that could be “automated”. That number can be misleading because it treats all work as equivalent. I prefer an autonomy ladder.
Level 1 — Retrieve: find and assemble trusted information. Level 2 — Summarise: compress the information into a usable brief. Level 3 — Recommend: propose a classification, priority or next action with evidence. Level 4 — Prepare: create the transaction, communication or change but require approval. Level 5 — Execute: take an approved class of action within defined boundaries. Level 6 — Decide: exercise material decision authority.
Most organisations can create significant value in Levels 1 to 4 without taking on the governance burden of Level 6. That is not a lack of ambition. It is good sequencing. The organisation learns about data quality, exceptions, monitoring and user trust while keeping irreversible decisions under clear human authority.
The control plane belongs in the architecture
Governance is often treated as a committee that reviews an AI solution after the interesting design work has happened. That is too late. Logging, tool permissions, identity, data boundaries, model selection, prompt and policy versioning, confidence thresholds, evaluation, exception routing and human override are part of the solution architecture.
The NIST AI Risk Management Framework and its Generative AI Profile are useful here because they frame AI risk across governance, mapping, measurement and management rather than as a one-time compliance check. ISO/IEC 42001 similarly treats responsible AI as a management system that has to be established, maintained and improved. The practical implication is simple: the control environment needs an owner, operating rhythm and evidence trail.
For agentic systems, I would add one more requirement: every material tool invocation should be attributable. If an agent updates a record, launches a workflow or generates a customer communication, the organisation should be able to reconstruct the triggering event, context, policy version, model output, approval state and resulting action.
Where I would start
The early portfolio should contain work that is frequent, coordination-heavy and easy to verify. Case summarisation, knowledge retrieval, document classification, duplicate detection, workload triage, test generation, code analysis and structured migration support are strong candidates because a person or deterministic control can inspect the output before consequence increases.
I would avoid starting with the most politically visible decision in the organisation. A first agent should earn trust by removing friction, not demand trust by asking for broad authority.
Five questions for the executive sponsor
- What coordination cost are we trying to remove?
- Which system is authoritative for every important fact the agent will use?
- Where does machine authority stop, and who owns the decision after that point?
- How will we evaluate the agent on real work, not just demonstration prompts?
- Can we reconstruct why an action occurred six months later?
Those questions are less exciting than a live agent demo. They are also much closer to whether the investment will survive contact with a real operating environment.
Where the economics actually come from
The most useful AI business cases I have seen do not start with headcount removal. They start with the cost of coordination. A service may have capable people and adequate systems yet still perform poorly because work repeatedly waits for someone to find context, reconcile two sources, determine ownership, request missing information, prepare a summary or hand the case to the next team. None of those activities is individually dramatic. Collectively they create queue age, rework and management overhead.
An agent can create value when it shortens that waiting time without weakening the control environment. That means measuring the baseline before automating: elapsed time between stages, touches per case, percentage of work returned for missing information, avoidable escalations, exception rate and the amount of senior staff time consumed by routine coordination. If those measures do not move, the AI initiative may be interesting but it is not transforming the operation.
Use progressive autonomy, not a binary decision
I prefer to design autonomy as a ladder. At the first level the system retrieves and summarises evidence. At the second it recommends an action. At the third it prepares the action but waits for approval. Only after evidence shows stable quality should it execute a bounded action itself. The progression should be based on observed error rates, consequence and reversibility—not enthusiasm for the technology.
That distinction matters in public services, finance, safety and other high-consequence environments. A system that automatically assembles the evidence for an eligibility assessment can save significant time while leaving statutory accountability with the officer. That may produce more real value than automating the final decision.
A practical executive test
Before approving an agentic use case, I would ask for five things on one page: the operational baseline; the authority the agent is being given; the data sources it may use; the failure and escalation paths; and the evidence that will be retained for audit. I would also ask what happens when the model, a source system or an integration is unavailable. If the proposal cannot answer those questions clearly, it is still a demonstration rather than an operating capability.
