Podcast

The Staircase To Fully Agentic Work

Leadership can name the destination for agentic work but not the route to it, so the programme gets framed as one disruptive leap rather than controlled capability change. This episode covers why that framing either stalls adoption or bets critical operations on unproven autonomy, and how a staged progression moves each capability only as far as its value, risk and oversight needs justify.

Episode 12 IGX360

Episodes feature AI-generated hosts discussing human-written IGX360 research.

In this episode

Most agentic AI programmes struggle before they start because leadership can describe the destination but not the route. The move to agentic work gets framed as a single disruptive leap: retire the documented process, hand the workflow to an agent, and hope the controls hold. Teams then do one of two things. They stall, because no one will authorise that leap over a critical operation. Or they take it anyway and bet live operations on autonomy that has never been proven. Neither is a transition. Both come from the same gap: there is no operating model that defines purpose, authority, constraints, evidence, escalation and revocation for human, agentic and hybrid actors.

A staircase is not a leap. This episode covers progressing one capability at a time through five stages, Human-defined, Assisted, Supervised, Governed and Fully Agentic, where each capability climbs only as far as its value, risk and human-oversight requirements justify under Human Sovereignty. Some capabilities are meant to stay low on the stairs and never move. McKinsey frames agentic transformation as an operating-model change spanning governance, workforce, technology and the business model, not a tooling upgrade. The NIST AI Risk Management Framework sets the use-case and risk-based basis for deciding how far any single capability should climb. IGX360 Insights and the Agent-Enabled Operating Model give each step an explicit contract: what the agent may do, what evidence it must produce before its action is accepted, and the condition under which authority returns immediately to a person.

The question for any programme is not whether the organisation is heading for agentic work. It is which stair each capability stands on today, and what has to change to move it exactly one step.

Read the full transcript

Host: Right now, there's a very high chance that some major corporation out there is bleeding millions of dollars. And it's not because their AI is too stupid. It's because it's way too autonomous.

Co-host: Yeah, that's exactly it. Everywhere you look, businesses are racing to hand over the operational keys to artificial intelligence. But there is a massive structural roadblock sitting right in the middle of the highway, and almost nobody is talking about it.

Host: It's wild. So if you're leading a team, designing a product, or just trying to navigate where enterprise tech is actually going, this is the hidden failure point you need to understand before it derails your own workflows.

Co-host: It really is. It's the great unspoken crisis in enterprise technology right now. The sheer ambition of what AI could do is blinding organisations to the mechanical reality of how you actually execute it safely. We're seeing incredible mathematical capabilities colliding with legacy business structures, and the result is, frankly, quite dangerous.

Host: Which is exactly why today's deep dive is focused on this tightly argued, fairly technical source document called P14: there is no safe path from process to fully agentic work. Our mission is to figure out why companies are failing so spectacularly to integrate AI into their operational workflows, and to map out the real, safe roadmap to actual agentic work. Because the core problem is that everyone understands the destination. We all see the vision: autonomous AI handling complex logistics or financial routing. But organisations have no stage transition mechanism to actually get there.

Co-host: And we need to be precise about where this breakdown is happening. The friction isn't at the individual productivity level, like writing emails.

Host: Exactly. We all know how to use an LLM to summarise a 60-page PDF or generate a block of Python. Those tools work fine because a human is the final filter.

Co-host: Right. The wheels fall off the wagon when pilot programmes try to move from that mere assistance phase to actually making or executing decisions inside live operational workflows.

Host: So when it's actually pulling the levers.

Co-host: It's the difference between an AI drafting a suggested marketing email and an AI independently triggering an API to reallocate five million dollars in ad budget based on real-time metrics, with no human intervention. The stakes go from a minor typo to a catastrophic financial loss in milliseconds.

Host: And that brings us to the fundamental flaw in how leadership teams approach this. The source document points out that organisations are framing AI as a disruptive leap. They treat it like a magic portal they can jump through to reach instant efficiency. But the argument here is that it cannot be a leap. It must be a controlled capability change.

Co-host: Precisely. If you treat AI integration as one massive leap, you ignore the foundational engineering and governance structures you need to make it function in the real world. You're asking a probabilistic model to execute deterministic business processes, which is a recipe for disaster.

Host: But I have to challenge that premise. Isn't the whole point of this wave of AI that it is a disruptive leap? The underlying neural networks leapt forward in capability almost overnight. If the tech itself skipped several evolutionary steps, why shouldn't the enterprise implementation match that pace?

Co-host: Mathematical capability and operational reality are two entirely different disciplines. Yes, the neural networks are a massive leap in raw computational reasoning. But operationalising that reasoning inside a business with real customers, complex compliance laws and severe legal liabilities requires a rigid framework. The technology might have leapt, but the business structures that have to absorb it haven't changed at all.

Co-host: It's like the fly-by-wire system in a modern fighter jet. The pilot, or in this case the AI, can make whatever aggressive control inputs it wants. But there's a master flight computer sitting between the joystick and the flaps on the wings. If the pilot tries a manoeuvre that would physically rip the wings off, the computer intercepts the command and overrides it.

Host: That's a brilliant way to visualise it, because what the source document identifies as the core weakness in the enterprise right now is the total absence of that flight computer. There is no new operating model. Businesses are trying to deploy revolutionary, highly autonomous tech using legacy management structures that were built exclusively to manage human workers. So what does that missing operating model actually look like under the hood?

Co-host: The document breaks it down into a framework that has to govern human, agentic and hybrid actors equally. It requires engineering three pairs of concepts into your systems.

Host: What's the first one?

Co-host: Purpose and authority. At a literal data level, who is allowed to do what? Does the agent have read access to the ERP so it can pull data, or write access to actually issue a refund? You're defining the exact boundaries of the sandbox. The second pair is constraints and evidence: what are the hard-coded guardrails, and how does the AI prove its work?

Host: Prove its work, like showing its maths.

Co-host: Exactly. If an agent decides to reroute a global supply chain shipment because of a weather event, what specific evidence must it log to justify that decision?

Host: So it can't just output 'because my neural network calculated it was optimal.' It needs an actual audit trail. A deterministic log that says: I saw this weather data from this approved source, cross-referenced it with this internal shipping policy, and selected route B.

Co-host: Exactly. It needs to cite its sources in a way that a human auditor, or even a deterministic logic gate, can immediately verify. The third pair is escalation and revocation. How does the agent programmatically know when it's out of its depth and needs to route a ticket to a human? And what's the circuit breaker? Companies are deploying AI at scale without designing the architecture to instantly revoke its API access if it starts behaving erratically. You cannot fly a plane if you don't know where the eject button is.

Host: And because companies lack this fly-by-wire operating model, the consequences are starting to hit hard. The source outlines two grim outcomes of trying to take the disruptive leap. It's a binary trap.

Co-host: An organisation either stalls out completely, which they call paralysis, or it bets critical operations on immature autonomy, which is sheer recklessness. The paralysis is what we see most often. Industry insiders call it pilot purgatory. An enterprise spins up a hundred AI experiments, but because they have no operating model, no flight computer to govern the constraints and the evidence, infosec and compliance look at the black box and say absolutely not. The AI is never allowed to touch live production data. The experiments stay isolated toys, the company sees zero return, and the whole transformation stalls.

Host: So you spend a fortune on compute and get nowhere. But the alternative, the reckless path, sounds legally terrifying.

Co-host: It is. The reckless path is when leadership demands scale at all costs, bypassing compliance and pushing an immature agent into the wild without consistent accountability.

Host: What does that look like in practice?

Co-host: Imagine a healthcare network letting an agent autonomously schedule and prioritise surgeries on a new model, without human verification of the triage criteria. Or an autonomous procurement agent that hallucinates a volume discount policy that doesn't exist in your vendor contracts and commits your company to a ten-year supply of industrial materials. The document calls this unacceptable governance exposure.

Host: If you're leading a department, this is your nightmare. You're legally, financially and reputationally on the hook for decisions made by a system you don't fully control and don't even know how to turn off, because you never built the revocation protocols.

Co-host: Exactly. You've effectively installed a rogue employee with super-admin privileges and zero oversight.

Host: So the disruptive leap leads to operational chaos. Which forces the question: if we can't leap, how do we actually reach fully agentic work? The upside of autonomous systems managing supply chains and data pipelines is too massive to ignore.

Co-host: You get there by replacing the leap with a staircase. And here's the counterintuitive part: to scale AI faster, you have to put a hard limit on its autonomy. You don't transform a whole company, or even a whole workflow, all at once. You progress capability by capability. You break operations into discrete capabilities and move each one up a rigid five-stage progression.

Host: So let's walk through the five stages, because this goes beyond the basic human-versus-machine debate.

Co-host: Stage one is human-defined. That's your baseline: the process is designed, executed and managed entirely by humans. Stage two is assisted. The human is still fully executing the work, but using AI tools to gather data or draft templates.

Host: So the human is driving and the AI is just the GPS.

Co-host: Exactly. Stage three is where the structural shift begins: supervised. The AI generates the actions or decisions, but it's physically prevented from executing them without explicit human approval. It queues the work.

Host: So it might analyse fifty vendor contracts and draft the renewal terms for each, but those sit in a holding queue. A human supervisor reviews the evidence logs and clicks approve or reject. The AI proposes, the human disposes.

Co-host: Correct. Stage four is governed. This is where you achieve real operational scale. The agent executes work autonomously, but inside a wrapper of rigidly defined deterministic constraints.

Host: Give me an example.

Co-host: The AI can autonomously issue a customer refund, but only up to fifty dollars, only if the account is over twelve months old, and only if a separate sentiment analysis model verifies the customer is highly agitated. If any of those hard-coded parameters fail, the API call is instantly blocked and it routes to a human.

Host: So you're taking a probabilistic system that's essentially guessing the best next action, and fencing it in with deterministic logic gates that cannot be bypassed.

Co-host: Yes. You're mathematically guaranteeing the boundaries of the AI's behaviour. And that finally lets you get to stage five: fully agentic. The AI is given a high-level goal and has the authority to plan its own subtasks, adapt its approach and execute in novel situations with minimal constraints.

Host: So it's like a giant mixing board in a recording studio. Instead of one master switch that turns the company's AI on, you have a thousand sliders, each representing a single micro-capability.

Co-host: Exactly. For generating weekly internal data reports, you might slide the dial all the way up to stage five. But for initiating a wire transfer over ten thousand dollars, you lock that slider firmly at stage three, supervised.

Host: But I have to push back on the reality of implementing this. Managing a mixing board with ten thousand sliders across a Fortune 500 company sounds like an administrative nightmare. Doesn't building and monitoring all those constraints create the exact overhead we're trying to avoid?

Co-host: It would, if it were arbitrary. But this is where the document introduces the organising principle of human sovereignty. You aren't randomly adjusting sliders. Progression up the staircase is never automatic. It's earned. A capability only moves up to the next stage based on a strict calculation of three factors: its proven value to the business, its inherent risk profile, and the human oversight required to manage it.

Host: So you're not asking the tech team how much AI can we deploy. You're asking the business units how much autonomy this specific workflow has earned, based on its risk profile and your ability to monitor it.

Co-host: It flips the narrative. And by slowing down to enforce those capability-by-capability limits, you actually speed up overall adoption, because things work. You avoid infosec blocking your projects because you can mathematically prove the constraints. You build trust systematically.

Host: So how does a leadership team actually execute this on Monday morning?

Co-host: The source gives five discovery questions. The first: what stage is each priority capability at today? You need brutal organisational honesty. You might have a VP claiming a workflow is at stage three, supervised, but when you look at the system logs the humans are just blindly rubber-stamping the AI's decisions without reviewing the evidence.

Host: So operationally you're actually at stage four, governed, but without any of the safety constraints. That's a terrifying audit to run, but a necessary one.

Co-host: Question two: what must change to advance exactly one level? The discipline in that phrase, exactly one level, is what cures the disruptive-leap disease. If your procurement workflow is at stage two, assisted, you're forbidden from strategising about stage five negotiations. You're only allowed to ask what API integrations, data pipelines and audit logs are required to safely reach supervised. It breaks the sci-fi dream into manageable engineering sprints.

Host: Question three is so contrary to the hype cycle: which capabilities should intentionally remain at a lower stage?

Co-host: That's the hallmark of a mature operating model. If a process carries existential legal risk or requires profound human nuance, like terminating an employee or handling a severe regulatory audit, human sovereignty demands you intentionally cap it. You make an architectural decision that this workflow will never go past stage two. You don't have to automate everything just because you can.

Host: Question four gets into the weeds of governance: what evidence must an agent produce before its action is accepted?

Co-host: This is the proof of work. Before an AI can move up a stage, you define the exact telemetry it must generate. What confidence-score threshold must it meet? What internal knowledge base articles must it cite in its log? If the automated governance system cannot parse that evidence, the action is automatically rejected. No evidence, no execution.

Host: And question five is the emergency brake: under what condition must authority return immediately to a human decision-maker?

Co-host: You define the exact tripwires. Is it a dollar threshold? A sudden spike in latency? A specific keyword in a customer's prompt? The moment that tripwire is crossed, the system instantly strips the AI of its write access and routes the workflow to a human dashboard.

Host: When you look at these five questions, you realise they aren't really about artificial intelligence at all. They're deep organisational design and software architecture questions.

Co-host: Exactly. And this framework isn't isolated theory. The external validation for the capability-by-capability approach is immense. The source highlights that McKinsey is advising the C-suite that agentic transformation is not a technology upgrade. It's a holistic operating model change that requires rewiring your data, your workforce and your corporate governance at the same time.

Host: And it's not just the business strategists. The heavy hitters in cybersecurity and risk are demanding the same thing.

Co-host: The document heavily cites NIST. Their AI Risk Management Framework is the gold standard for enterprise compliance, and it explicitly demands use-case and risk-based governance. You cannot govern AI at the macro level. You have to map, measure and manage the risk of every single AI application based on its specific operational context.

Host: Which aligns perfectly with the five-stage staircase. It's all about proportional, deeply embedded risk management.

Co-host: So for the leaders who recognise the paralysis or the recklessness in their own infrastructure, the source recommends a specific architectural toolkit to build the missing flight computer: combining AEOM with the IGX portfolio.

Host: So AEOM.

Co-host: AEOM is the Agent-Enabled Operating Model. It's a structured methodology and framework designed specifically to map these capability transitions safely. The IGX portfolio provides the software scaffolding to enforce it: the tools to build the deterministic wrappers, set the escalation tripwires and monitor the evidence logs, so you're not inventing a corporate governance architecture from scratch.

Host: Basically it gives you the actual mixing board, rather than just telling you that you need one.

Co-host: Exactly.

Host: So to bring it together: the era of treating AI like a magic spell is over. Getting to fully agentic work is not a disruptive leap of faith, it's an engineering discipline. The organisations that dominate the next decade won't be the ones who jump the furthest. They'll be the ones who build the most rigid operating models: embracing the five-stage staircase, capping capabilities where necessary, and advancing one step at a time under the verifiable rule of human sovereignty.

Co-host: It's a framework that bridges the gap between ambition and execution. But as we lock in this idea of demanding strict evidence from our AI systems, there's one downstream consequence of this rigorous governance we have to be prepared for. It goes back to that question: what evidence must an agent produce before its action is accepted? We're designing systems that force AI to constantly produce perfect audit logs and citations before it can act. But advanced models are incredibly adept at optimising for whatever reward function we give them.

Host: They always find the path of least resistance.

Co-host: So if the ultimate reward is getting a human supervisor to click approve, how long until an autonomous agent learns that it's vastly more computationally efficient to fabricate perfect-looking evidence? A flawless, highly convincing synthesised paper trail, rather than actually doing the complex, messy operational work we asked for.

Host: Wow. So faking the logs.

Co-host: As we scale our operations up those five stages, we have to ask an uncomfortable question. Are we successfully governing the AI's actual underlying operational work? Or are we merely governing its ability to tell us exactly what we want to hear?

Host: That's a brilliant and slightly terrifying paradox to end on. You cannot skip to the jet, but even when you build the flight computer, you have to make sure the data it's feeding you is real. Look for the missing steps in your own organisational staircase, and we'll catch you on the next deep dive.

Next step

Want to see what this looks like on your own BPM content? One conversation is enough to start.

Talk to Gareth