Podcast

Why AI Pilots Lack Operating Value

Billions go into enterprise AI assistants and agents while operating value stays flat. This episode covers why pilots bolted onto an unchanged operating model never scale, and what defining purpose, authority, constraints, evidence, escalation and revocation for every actor does to close the gap.

Episode 12 IGX360 14:33

Episodes feature AI-generated hosts discussing human-written IGX360 research.

In this episode

Billions are going into enterprise AI assistants and agents, and the needle on operating value has barely moved. The reason is not model quality. It is that most pilots run in isolated functions with the underlying work left untouched, so the technology sits beside the operating model rather than changing it. The gap becomes obvious the moment a pilot moves from helping one person to making or executing decisions inside a live workflow, where nobody has defined the agent’s purpose, authority, constraints, evidence, escalation or revocation.

This episode works through P11 with its McKinsey and NIST anchors: experimentation is everywhere, scaling reaches one or two functions, and fragmented pilots create governance exposure that ordinary shadow IT never did. The fix is deliberate, not opportunistic. Map each capability, redesign the work, and let autonomy advance only to the level its value, risk and oversight needs justify, with the constraints and escalation triggers built into the workflow before an agent runs. iGrafx models where an agent acts and where a human gate belongs. IGX360 Insights monitors execution so the evidence holds up. The closing question for any leader: can you name one pilot that changed an end-to-end workflow, or only ones that sped up a task?

Read the full transcript

Host: You've probably noticed that everywhere you look right now, companies are launching these flashy new AI assistants.

Co-host: Oh, absolutely everywhere. You can't escape it.

Host: Right. But here's the thing. Why are so few of them actually transforming bottom line business value?

Co-host: Yeah, that is the million dollar question. Or well, maybe the billion dollar question at this point.

Host: Seriously. The billions are being spent, but the actual needle on the company's operating value, it's barely moving.

Co-host: So welcome to today's deep dive. We are unpacking this massive gap between all the hype of enterprise AI experimentation and actual operating value.

Host: And it really is a huge gap. We're seeing organisations just throwing capital at these agents to test them out. But when you look at the structural economics, like the actual operational leverage, it's just flat.

Co-host: Yeah, it's completely stalled. So to figure out why, we are looking at a highly specific strategy document today. It's titled P11: AI Pilots Do Not Become Operating Value.

Host: A punchy title, right. But it's actually a fantastic breakdown. It uses a SPIN framework: situation, problem, implication, need, payoff.

Co-host: Right, and it doesn't just make empty claims. It backs everything up with external validation from McKinsey and NIST, the National Institute of Standards and Technology.

Host: Exactly. It's very grounded. So our mission for this deep dive is to figure out why this tech is failing to scale at the enterprise level right now, and what this new framework they call Human Sovereignty can do to fix it.

Co-host: Let's just jump right into it.

Host: We should probably start by diagnosing the exact deployment flaw, right? Like the situation and the problem.

Co-host: Yeah, let's start there. How are companies actually deploying these things?

Host: The fundamental problem is that teams are experimenting with these assistants and agents in totally isolated functions. And crucially, they are doing this without redesigning the underlying work.

Co-host: Meaning they aren't changing the day-to-day processes at all.

Host: Exactly. The technology is just sitting completely beside the operating model. It's not changing it. And this flaw becomes glaringly obvious when these pilots try to move from simply assisting a single person to actually making and executing decisions inside real operational workflows.

Co-host: Okay, let's unpack this, because the document makes a really big deal out of the sitting beside the operating model thing.

Host: That's the core issue, yeah.

Co-host: It makes me think of an analogy. It's like buying a state-of-the-art jet engine, right?

Host: Okay, I'm with you.

Co-host: And you take this incredible million dollar jet engine, but you just strap it to a wooden, horse-drawn carriage.

Host: Oh wow, yeah, that's a perfect visual.

Co-host: Right, like you didn't change the carriage's design, you didn't upgrade the wheels or the steering, you just bolted on this massive new power source.

Host: But let me ask you this. Is the problem really the AI tech itself? Or is it that we are treating these new agentic and hybrid actors like regular software without defining the rules of engagement?

Co-host: It is 100% the latter. The technology is the jet and it works. The underlying weakness is the absolute absence of an operating model for that engine.

Host: So we just don't have a framework for how they should operate.

Co-host: Right. The document explicitly lists six critical definitions that are just entirely missing when companies deploy these sidecar AI agents.

Host: Wait, six definitions. What are they?

Co-host: Okay, so they are purpose, authority, constraints, evidence, escalation, and revocation.

Host: That's a lot. Let's break those down a bit. Purpose and authority seem pretty straightforward, right?

Co-host: Yeah, purpose is just the objective, like reduce shipping costs. Authority is what the AI is actually allowed to execute independently.

Host: Like approving a refund up to 50 bucks or whatever.

Co-host: Exactly. But then you get into constraints, which are the hard boundaries. It absolutely cannot change a vendor's core contract.

Host: Right, the guardrails.

Co-host: And the really critical ones for oversight are evidence, escalation, and revocation. Evidence means the AI has to prove its work before taking action.

Host: Like showing its math.

Co-host: Right. Escalation is knowing exactly when to flag a human because it hit an edge case. And revocation is the kill switch. How do you instantly pull its access if things go wrong?

Host: So if you're listening to this and your workplace is rolling out some new AI tool, you really have to look around and ask, are we actually redesigning the work with these six definitions? Or did we just strap a jet engine to the side of our old carriage processes?

Co-host: And if you're just strapping it on, you're heading for disaster. Which, if we connect this to the bigger picture, leads directly to this massive enterprise-wide scaling crisis.

Host: The fallout, basically.

Co-host: Yeah, the implication phase of the document. Because these pilots are just sitting beside the operating model, they multiply rapidly, but governance completely fragments.

Host: Meaning everyone is doing their own thing.

Co-host: Right, and those promised EBIT numbers, you know, earnings before interest and taxes, the big productivity impacts, they just remain completely elusive.

Host: Because they're stuck in isolation.

Co-host: Exactly. At an enterprise scale, these experiments either stay isolated or they try to scale without consistent accountability. The value is totally unclear and it creates this unacceptable governance exposure.

Host: Yeah, and the document brings in some heavy hitters to prove this.

Co-host: It does. It cites McKinsey's The State of AI 2025 report, which is great external validation. McKinsey highlights that sure, there is widespread experimentation everywhere, but the actual scaling is incredibly narrow.

Host: Narrow, like they only use it in one department.

Co-host: Usually just one or two functions, yeah. Because without an operating model, they can't safely expand it.

Host: Well, wait, let me push back on this a little bit, especially this idea of unacceptable governance exposure. If every department is building their own isolated pilot to fix a siloed problem, aren't we just talking about the old shadow IT problem?

Co-host: Yeah, I mean, 10 years ago, marketing was using unapproved Dropbox accounts and HR was using some random survey software. It accelerated fragmentation, sure, but the company didn't collapse.

Host: How is an organisation possibly going to manage risk when everyone is running these new shadow AI experiments? Is it really that different?

Co-host: It is fundamentally different, and that's a great question, but you have to look at what the software is actually doing. Traditional shadow IT involved humans using unauthorised tools.

Host: Right, the human was still the one acting.

Co-host: Exactly. Shadow AI involves unauthorised entities making actual operational decisions. A rogue Dropbox account might leak a file, which is bad. But an unapproved, unconstrained AI procurement agent might autonomously hallucinate and sign a $10 million vendor contract.

Host: Oh, yeah, okay, that is a wildly different risk profile.

Co-host: Because the authority and constraints weren't defined. So you just have these fragmented agents making decisions based on invisible logic across the whole company.

Host: It's a governance nightmare. So how do we fix it? Because the source doesn't just complain about the problem. It introduces a need-payoff strategy, right?

Co-host: Yes, it introduces a major philosophical shift to solve this exact lack of accountability. Here's where it gets really interesting. The document calls this blueprint Human Sovereignty.

Host: Human Sovereignty, yes. I love that term.

Co-host: Basically, it means you have to map out your capabilities and redesign the work so that AI autonomy progresses deliberately, not just opportunistically because someone thought of a cool prompt.

Host: Right, you don't just deploy it because you can.

Co-host: Every single capability has to progress to a level of autonomy that is strictly justified by its value, its risk, and its human oversight requirements.

Host: You're actually putting the human at the centre of the sovereignty model.

Co-host: Exactly. And the indicated benefits are massive. You get fewer disconnected pilots. You get actual workflow-level value instead of just task-level shortcuts, and you establish really clear boundaries for authority.

Host: Which gives you credible AI value cases finally.

Co-host: Right, backed by continuous evidence and oversight. And to actually do this, they bring in that second external validation, the NIST AI Risk Management Framework.

Host: NIST, the National Institute of Standards and Technology. How does that fit in?

Co-host: Well, NIST provides this incredibly rigorous use-case and risk-based framework. It gives leaders a way to govern, map, measure, and manage AI risk across the entire lifecycle of the system.

Host: So it's basically the playbook for figuring out what risk tier a specific AI agent falls into.

Co-host: Exactly. It gives you the structure to enforce human oversight.

Host: Okay, but in practical terms, how do we actually enforce this sovereignty when these agents are operating at a speed humans just can't match?

Co-host: What do you mean?

Host: Well, the source mentions a controlled progression from assistance to autonomy, right? But if an AI is optimising supply chain routes in real time, crunching thousands of variables a second, how is a human supposed to stay sovereign over that? We can't think that fast.

Co-host: Right, that's a really common misconception. You don't enforce Human Sovereignty by manually reviewing every single micro decision in real time.

Host: Because then you just lose all the speed benefits of the AI.

Co-host: Exactly. You enforce it by designing the architecture of the workflow itself. You build the constraints and the escalation triggers into the system before the AI is even allowed to operate.

Host: Oh, I see. So you build a really strict maze and the AI just runs inside the maze.

Co-host: You dictate the rules of the environment. And to help leaders actually audit their own mazes, the source provides this diagnostic toolkit.

Host: Yeah, the discovery questions. This is where it shifts from theory to a really practical interrogation of your own business.

Co-host: Right. It gives leaders a way to expose the sidecar illusion in their own ranks.

Host: So what does this all mean? Well, let's actually run through these five discovery questions. And if you're listening to this, treat this as a thought exercise for your own company.

Co-host: Definitely. Question one is a gut check. Which AI pilots changed an end-to-end workflow?

Host: Not just a single task, but the whole workflow.

Co-host: Right. Question two, what operating-model constraint prevents scale?

Host: Meaning what old rule is stopping this new tech from working?

Co-host: Exactly. Question three asks, how are pilot outcomes compared across capabilities?

Host: Are we even measuring these things against each other?

Co-host: And then questions four and five are the really mechanical ones. Question four, what evidence must an agent produce before its action is accepted?

Host: Let me stop you right there, because question four is fascinating to me. Does this framework fundamentally shift the burden of proof onto the AI?

Co-host: How so?

Host: Well, normally we test software, we say it works, and then we just trust it to run in the background. But this question implies that we shouldn't just trust the AI based on its potential. It's no longer about what the AI can do, but what it can prove it did correctly in real time.

Co-host: This raises an important question, and you're spot on. It absolutely shifts the burden of proof. It demands a posture of continuous evidence. You aren't trusting the algorithm. You are trusting the evidence it produces at every step.

Host: Wow. Okay. And what was the fifth question?

Co-host: Question five is the kill switch. Under what condition must authority return immediately to a human decision-maker?

Host: So defining exactly when it gives up and asks for help.

Co-host: Right. And the document actually has a call to action here. It urges leaders to book a call and honestly evaluate why their current pilots have or have not changed end-to-end business performance.

Host: Which requires a lot of honesty.

Co-host: Oh, totally. But the source doesn't just leave you with questions. It provides structural solutions through a product route.

Host: Right, internal tools. Like the AEOM maturity model.

Co-host: The Agentic Enterprise Operating Model. That maturity model is the roadmap. It helps you assess exactly where your workflows are lacking.

Host: And it's supported by software like iGrafx.

Co-host: iGrafx does the process modelling. It acts as the digital architecture where you actually build those constraints we talked about. You map the exact nodes where an agent acts and where a human takes over.

Host: So it draws the maze.

Co-host: Exactly. And then IGX360 Insights layers on top of that.

Host: To do what?

Co-host: To monitor the workflow in real time. It ensures that the continuous evidence architecture is actually working. It's how you guarantee that these AI pilots finally transition into real, tangible operating value.

Host: So it's about taking the abstract idea of we need rules and actually hard coding it into your digital infrastructure.

Co-host: That is exactly it. So let's briefly recap what we've covered today, because this is a lot to digest.

Host: We started by looking at this massive problem, these isolated sidecar AI experiments that companies are doing.

Co-host: Right, strapping the jet engine to the carriage.

Host: Exactly. And we saw how that lack of an operating model just fragments governance and completely fails to deliver actual EBIT value.

Co-host: As proven by the McKinsey data.

Host: Right. Then we moved to the solution. Deliberately redesigning workflows under this framework of Human Sovereignty.

Co-host: Using the NIST AI Risk Management Framework to map out the risk.

Host: Exactly. Finally, we looked at how to interrogate those systems using those five discovery questions, and using tools like iGrafx to demand continuous evidence.

Co-host: It's a complete paradigm shift for enterprise AI.

Host: It really is. But before we go, I want to leave you, the listener, with a final provocative thought, something to just mull over that builds on these rules we just discussed.

Co-host: Oh, I'm curious where you're going with this.

Host: Well, think about question five, the revocation constraints. And question four, the strict evidence requirements.

Co-host: Right, the oversight mechanisms.

Host: Let's say a company does this perfectly. They build the ultimate Human Sovereignty architecture. The AI is constantly pausing to prove its work, or it's instantly escalating to a human the second a complex parameter is breached. If we perfectly design a system that demands human oversight for every critical juncture, does the ultimate bottleneck to enterprise scale just become the human decision-maker waiting for authority to return?

Co-host: Oh, wow. So you're saying the human becomes the weak link in the chain?

Host: Exactly. Are we intentionally designing a system where the AI's ultimate speed and scale are permanently capped by our own biological need for Human Sovereignty? I mean, if the AI is always waiting for us to verify its evidence, can we ever truly reach infinite scale?

Co-host: That is a fascinating tension. The machine is optimised, but it's tethered to our processing speed.

Host: Right, it's a structural paradox that every leader is going to have to navigate as these agents get smarter and faster.

Co-host: Thank you so much for joining us on today's deep dive. Keep asking the tough questions, keep challenging those models, and we will catch you next time.

Next step

Want to see what this looks like on your own BPM content? One conversation is enough to start.

Talk to Gareth