Podcast

Deciding Where AI Agents Should Act

Autonomy gets granted by enthusiasm, not capability, so the debate collapses into automate everything or block it. This episode covers the AEOM 1 to 5 maturity dial: a repeatable way to set a justified autonomy target for each capability, release low-risk value sooner, and keep a human gate where judgement and accountability belong.

Episode 13 IGX360 25:20

Episodes feature AI-generated hosts discussing human-written IGX360 research.

In this episode

Autonomy tends to get granted by enthusiasm rather than capability. Someone demonstrates a convincing agent, and the conversation splits into two camps: automate the whole function, or block it until the risk disappears. Neither camp is actually deciding anything. The organisation has no shared method for selecting an appropriate level of autonomy capability by capability, and the gap shows most clearly when a pilot stops assisting one person and starts making or executing decisions inside a live workflow.

This episode works through P12 with its NIST and EU AI Act anchors. Both point the same way: govern AI by context and risk, not with a single undifferentiated control. The AEOM 1 to 5 maturity dial puts that into practice. Each capability gets a justified target autonomy level, set against its value, its risk and the human oversight it needs, so low-risk work releases value sooner while high-judgement decisions keep a human gate. iGrafx models where an agent acts and where that gate belongs. IGX360 Insights monitors execution so the evidence behind each step holds up. The question for any leader: can you name the method you used to decide an agent’s autonomy, or was it just how convincing the demo felt?

Read the full transcript

Host: So picture this scenario for a second. You are sitting at your desk, it's Tuesday morning, and you're just reviewing your normal dashboard.

Co-host: Just a regular Tuesday.

Host: Right, exactly. And you've had this AI assistant integrated into your enterprise workspace for like the last six months or so. And up until now, it's basically been the ultimate quiet intern.

Co-host: Oh, I know the type. You mean summarising lengthy status meetings, that sort of thing.

Host: Yeah, exactly. Summarising meetings, maybe drafting supplier emails that you still manually approve, or pulling raw data into your spreadsheets. But then today, something shifts.

Co-host: Okay, I feel like this is where it goes off the rails.

Host: Well, a major weather event hits a key shipping lane. And before you even pour your morning coffee, you get a system notification. The AI hasn't just summarised the supply chain delay for you.

Co-host: Oh, wow. What did it do?

Host: It independently identified three alternative suppliers, negotiated a spot rate, updated your company's financial ledger, and rerouted $2 million of incoming inventory.

Co-host: Wait, 2 million?

Host: Just like that.

Co-host: Just like that.

Host: Inside your live operational workflow.

Co-host: I mean, it didn't flag you for permission or anything. It just executed the entire pivot completely autonomously.

Host: Right. And see, that is the exact inflection point. That moment where the machine transitions from just being a passive advisor to, well, an active agent. And it represents the specific threshold that is currently, honestly, paralysing business leaders globally right now.

Co-host: Because it's a massive leap.

Host: Huge.

Co-host: I mean, we are watching the transition from AI as just a tool that generates text to AI as an actual actor that manipulates enterprise infrastructure.

Host: And that scares people. Which is exactly why we are dedicating today's deep dive to a strategy document that kind of functions as a survival guide for that exact transition.

Co-host: It really is a survival guide. Yeah.

Host: The source text we're looking at today is titled P12: We cannot decide where agents should act.

Co-host: Great title, by the way. Gets right to the point.

Host: It really does. And so our mission today is to dissect exactly why organisations freeze up when they try to figure out how much autonomy to give AI. And we're going to map out the specific maturity framework this document actually proposes to fix that whole problem.

Co-host: Because it is a highly fixable problem, even if it doesn't feel like it right now.

Host: Exactly. So if the AI conversations in your boardroom currently feel like this exhausting, chaotic tug of war, we are going to provide you with the blueprint. We'll help you understand the structural failures that are causing that friction and give you the operational model to resolve it.

Co-host: Right. And to understand the solution, we really have to look closely at that structural failure you just mentioned.

Host: Let's get into it. What's actually failing?

Co-host: Well, the document diagnoses a very specific operational gap. Basically, organisations currently possess absolutely no shared, standardised method for selecting appropriate autonomy by capability.

Host: Meaning they just don't know how to measure who gets to do what.

Co-host: Exactly. What happens in practice is that you have these small AI pilots, you know, the tools summarising those meeting notes we talked about, and they naturally begin to drift toward execution.

Host: Right, creeping into the actual work.

Co-host: Yeah, they start interacting with live operational workflows. And when that line gets crossed, leadership suddenly realises, wait a minute, we have absolutely no framework to evaluate whether the AI should possess the authority to take that action.

Host: That total lack of a framework creates what the text calls this toxic binary trap.

Co-host: Oh, the binary trap is the worst part.

Host: It really is. Because there's no method to measure or assign autonomy, the internal debate just instantly polarises. It's all or nothing.

Co-host: Yeah, all or nothing.

Host: You get one camp that points to the incredible efficiency of rerouting that supply chain we mentioned and they demand: automate everything, look how fast this is. But then the risk and compliance camp points to that unapproved $2 million ledger update and they just scream: block it entirely.

Co-host: Which is a natural reaction when you don't have control.

Host: Sure, but it ends up meaning we are treating artificial intelligence like a standard light switch when, honestly, the technology fundamentally requires a dimmer.

Co-host: A dimmer switch. I love that analogy.

Host: It just makes zero sense to govern a really low stakes task like, I don't know, filtering spam vendor invoices with the exact same binary switch that we use for a mission critical decision, like reallocating millions in capital.

Co-host: It makes no sense at all. And that light switch analogy perfectly illustrates how businesses are applying these legacy software mentalities to what are actually probabilistic systems. But the document digs into the root cause of this binary trap. The panic doesn't actually stem from the technology itself.

Host: It doesn't.

Co-host: No, it stems from the total absence of an enterprise operating model.

Host: Okay, let's avoid the corporate jargon trap here, though. When the text says operating model, it's not just talking about some generic mission statement framed on the wall.

Co-host: Not at all. Think about it. We have centuries of established operational governance for humans. We have spending limits, localised escalation matrices, compliance audits.

Host: I have to get like four different signatures.

Co-host: Exactly. But for agentic actors, for AI, the enterprise architecture is essentially a blank slate right now.

Host: Wow, a blank slate.

Co-host: And that blank slate is the core vulnerability. The text defines this missing operating model by literally listing out its necessary components.

Host: Okay, what are we missing?

Co-host: Well, organisations currently lack enterprise definitions for purpose, authority, constraints, evidence, escalation, and revocation for these human, agentic, and hybrid actors.

Host: So if something goes wrong, we have no plan.

Co-host: Basically, if an AI agent hallucinates a supplier rate, there is literally no defined escalation path. Or if it begins executing unauthorised ledger updates, there is no standardised revocation protocol to instantly pull its access.

Host: It's just out there doing whatever.

Co-host: Yeah. And without those rigid enterprise definitions mapped out in advance, your only available responses are to carelessly flip the switch to on or completely panic and slam it to off.

Host: Okay, so if this total lack of a framework is causing all this paralysis, we really have to look at how much this fear is actually costing these companies.

Co-host: Oh, the cost is staggering. The text points out this fascinating dual consequence of the light switch mentality. By slamming the switch to off, out of pure fear, all this low risk value is being severely delayed.

Host: Just left on the table.

Co-host: Yeah, organisations are literally leaving thousands of hours of safe administrative automation behind because they just don't have a way to compartmentalise the risk.

Host: Right, and the stagnation of that low risk value, it's basically a massive hidden tax on enterprise productivity.

Co-host: Definitely. But the second consequence the document outlines is honestly far more insidious.

Host: Oh.

Co-host: Yeah, because while the safe, low risk capabilities are blocked by all this bureaucracy, high-risk deployments are simultaneously advancing without any adequate safeguards.

Host: Wait, I need to challenge the mechanics of that. If leadership is actively blocking AI because they're terrified of the risk, how on earth does a high-risk deployment bypass them? I mean, are we really saying companies are stifling safe innovation while accidentally maximising their risk exposure? That sounds like the worst of both worlds.

Co-host: It is the worst of both worlds. And that is exactly the paradox the document highlights when you get to enterprise scale.

Host: Okay, how does that happen?

Co-host: Well, when leadership is trapped in that polarised binary debate at the very top, they fail to provide concrete operational governance for everyone else. And nature abhors a vacuum.

Host: And so do software developers.

Co-host: Exactly. A single department, or honestly even just a single frustrated engineer, will simply go out and procure an AI agent to handle some complex, high-risk data.

Host: Because nobody explicitly gave them an operating model that said they couldn't.

Co-host: Yes. What often happens is a small pilot programme that was designed for a super-controlled, low-risk environment just gets quietly scaled up across the entire network without anyone auditing its new authority.

Host: Wow, so it's basically shadow IT, but instead of someone just downloading an unauthorised messaging app, they're plugging an autonomous decision engine directly into the company's nervous system.

Co-host: That is a terrifying but very accurate way to put it. And the source document explicitly outlines two massive enterprise level headaches this produces.

Host: What are they?

Co-host: First, you get unclear value. Because these deployments are isolated, ungoverned, and just hidden in the shadows, no one is rigorously tracking the return on investment.

Host: So you don't even know if it's saving money.

Co-host: Right. But second, and far more critical, you generate unacceptable governance exposure.

Host: Because it's operating in the dark.

Co-host: Exactly. You have probabilistic models making real operational decisions in a vacuum, completely devoid of those constraints or escalation paths we discussed earlier. I mean, that is the literal definition of an operational time bomb.

Host: Yeah, we clearly need that dimmer switch to replace the binary light switch.

Co-host: Desperately.

Host: Which introduces the core solution proposed by the document. The text calls it the AEOM 1 to 5 maturity dial.

Co-host: Yeah, let's break down that acronym first, because the terminology actually really matters here. It stands for Agentic Enterprise Operating Model, and that phrasing, Enterprise Operating Model, is really the key takeaway. It's not just a software manual. It's not just a software deployment guide. It is a fundamental restructuring of how work itself is governed.

Host: Right.

Co-host: The AEOM 1 to 5 dial is designed to set a justified target per capability. So you are not setting an autonomy level for the entire marketing department or the whole supply chain division.

Host: Okay, so it's much more granular.

Co-host: Exactly. You are deconstructing those massive divisions into micro capabilities and you're assigning a specific dial setting to each individual task.

Host: And the philosophy anchoring this entire 1 to 5 dial is something that the text calls Human Sovereignty.

Co-host: I love that term.

Host: It's powerful, right?

Co-host: Yeah.

Host: It dictates that the human remains the absolute sovereign ruler of the workflow. Period. Any autonomy granted to the machine is conditional, it's temporary, and it's aggressively monitored.

Co-host: The text states that a capability only progresses up the 1 to 5 dial when it's justified by its value, its risk profile, and its human oversight requirements.

Host: Right. And to really contextualise how that works, consider how compartmentalised governance limits systemic risk in other engineering disciplines.

Co-host: Oh, like circuit breakers in a commercial power grid.

Host: Exactly. Like if you plug a faulty piece of equipment into your kitchen outlet, a localised level 1 or level 2 breaker trips.

Co-host: Right, the lights go out in the kitchen.

Host: Yeah, it isolates the failure. It doesn't shut down the entire neighbourhood's power grid.

Co-host: So under this AEOM dial, giving an AI level 5 autonomy to, say, independently match invoice numbers, that doesn't grant it grid level access to authorise the final $1 million wire transfer.

Host: No, absolutely not. The risk is physically compartmentalised by the capability's specific setting on the dial.

Co-host: The circuit breaker analogy captures the mechanism perfectly. And the document lists some highly specific benefits to adopting this model.

Host: What's the biggest one?

Co-host: Well, first, it forces faster agreement on appropriate autonomy. You basically strip all the emotion out of the boardroom.

Host: That sounds like a miracle.

Co-host: Right. You aren't sitting there debating the existential threat of AI taking over the world. You are simply agreeing that, hey, cross-referencing vendor addresses is a low-risk capability suitable for level 2 autonomy.

Host: And everyone can agree on that.

Co-host: Exactly. This allows the enterprise to rapidly release value in lower risk work without overexposing their high judgment capabilities.

Host: So you basically unblock the administrative backlog immediately.

Co-host: Yes. Furthermore, the text argues this establishes clear autonomy and authority boundaries, which leads directly to more credible AI value cases.

Host: Which CTOs love.

Co-host: They need it. Because if I'm a CTO, I can now definitively prove the ROI of an AI agent operating at level 3 on a specific data entry task.

Host: That's way better than just throwing $1 million at some vague enterprise AI initiative and just hoping productivity magically increases.

Co-host: Exactly. But the final benefits outlined are perhaps most vital for actual compliance.

Host: Okay, let's hear them.

Co-host: The model mandates continuous evidence and oversight, and it enforces a controlled progression from assistance to autonomy.

Host: Meaning you can't just skip steps.

Co-host: Right. You do not leap from drafting emails to rerouting entire supply chains. You turn the dial incrementally, much like getting a learner's permit before a driver's licence.

Host: Oh, I like that.

Co-host: Yeah, you require the AI to produce irrefutable evidence of its reliability at every single stage, in safe neighbourhoods, before it is ever granted further authority on the highway.

Host: Okay, I see the theory there and it sounds great. But if I'm an operations manager on the ground, implementing a five-tier dial sounds like a massive bureaucratic bottleneck.

Co-host: It can sound daunting.

Host: Like, how do I physically know where my specific supply chain workflow even belongs on that dial?

Co-host: Well, the document addresses this by moving away from just theory. They introduce five specific discovery questions.

Host: Oh, okay. The intake process.

Co-host: Exactly. This is the actual intake process.

Host: Let's dissect these, because this is where the operating model actually gets built.

Co-host: Let's do it.

Host: So question one asks: which capabilities are high volume and low judgment?

Co-host: High volume, low judgment.

Host: Okay, so this first question basically maps your operational topography.

Co-host: Exactly. High volume capabilities consume massive amounts of expensive human capital, but low judgment means the task operates on really deterministic rules.

Host: Right.

Co-host: It does not require nuanced human empathy.

Host: Or complex ethical trade-offs.

Co-host: No, none of that. And identifying these tasks first immediately reveals your prime candidates for high autonomy deployment.

Host: So let's apply this to the supply chain scenario. Taking incoming shipping manifests and manually checking if the container weights match the contracted manifest.

Co-host: High volume, low judgment. It's just pure data verification.

Host: Crank that dial up to level 4, right?

Co-host: But then we hit the second discovery question, which asks: which decisions demand human review or accountability?

Host: And this question is vital because it establishes the absolute ceiling for the AI's autonomy.

Co-host: Even if it's high volume.

Host: Yes. Even if a task is high volume, if the consequence of an error is severe, the capability cannot advance to full autonomy.

Co-host: But who defines what high judgment is, though? Like deciding to drop a supplier entirely because their last three shipments were late. That is a decision that impacts long-term vendor relationships.

Host: It has potential contract penalties.

Co-host: Huge implications.

Host: That demands human accountability.

Co-host: The AI can analyse the late shipments all day and maybe draft a recommendation, but a human must be the one to actually authorise that contract termination.

Host: Exactly. And just like that, we've mapped the safe valleys and the dangerous peaks of the workflow.

Co-host: Okay, so once the topography is mapped, the questions shift to the mechanics of governance.

Host: Yes. Question 3 asks: what evidence would justify moving a capability up one stage?

Co-host: And this is where we really need to get technical.

Host: Yeah, this is the heavy lift. Because the text shifts the conversation from these vague, subjective feelings of, oh, I trust the AI, to concrete operational metrics.

Co-host: So if I'm auditing a probabilistic black box like a large language model, what does evidence actually look like in an enterprise system?

Host: Well, it looks like rigorous statistical proof of performance over a set amount of time.

Co-host: Not just a demo.

Host: No, definitely not a demo. Evidence is an accuracy rate of 99.98% over 50,000 parallel transactions.

Co-host: Oh, okay.

Host: It is a zero variance audit trail that is verified by a secondary deterministic software system. The business basically has to rigidly define the mathematical threshold for success before the capability is allowed to move from level 2 to level 3 on the dial.

Co-host: And that structural requirement leads directly into question 4, which asks: what evidence must an agent produce before its action is accepted?

Host: Yes, and this is a crucial distinction.

Co-host: Right, because question 3 is about systemic promotion, moving up the dial over time. But question 4 is about the micro level, the transactional reality.

Host: Even if an AI is authorised to operate at level 4, what artefacts must it generate for every single action it takes? The agent must generate a deterministic receipt for a probabilistic action.

Co-host: A deterministic receipt. I like that.

Host: Yeah. So if the AI agent reroutes that delayed inventory we discussed earlier, it cannot simply update the ledger and move on.

Co-host: It has to show its work.

Host: It has to append the specific API logs it queried regarding the weather event. It needs to show the predictive confidence score of the new shipping route. And it has to link the exact policy document it cross-referenced to authorise the spot rate.

Co-host: It explicitly shows its computational work every single time.

Host: Yes, exactly. Which ensures that the system, regardless of its autonomy level, is fully transparent and auditable by a human at any given second.

Co-host: Fully auditable.

Host: Which brings us to the 5th and final discovery question. And arguably, this is the most critical one for enterprise safety: under what condition must authority return immediately to a human decision maker?

Co-host: This is the tripwire.

Host: The tripwire.

Co-host: This is the exact definition of that revocation protocol we noted was missing earlier.

Host: Meaning you have to define the edge cases before you ever even turn the dial.

Co-host: Before we even touch the dial. Right. So if the negotiated shipping spot rate exceeds the historical average by more than 15%, boom, the circuit breaker trips and authority returns to a human.

Host: Instantly.

Co-host: If the vendor API returns a degraded confidence score, the breaker trips. You engineer the exact operational thresholds that instantly strip the AI of its autonomy and enforce that Human Sovereignty we talked about.

Host: And by answering these five discovery questions, an organisation basically transforms the nebulous, terrifying concept of autonomous agents into a highly structured, mathematically governed operational matrix.

Co-host: You've effectively installed the dimmer switch.

Host: You've installed it, and you've engineered the precise constraints for its operation.

Co-host: But okay, let me play devil's advocate. The EU AI Act, for example, is notoriously dense and legally complex. If I'm a compliance officer, I'm going to look at this AEOM framework and ask: does a simple five-point dial actually satisfy a European regulator? Or is this just some corporate strategy team trying to sound compliant?

Host: That's a fair question.

Co-host: But the document addresses this scepticism, right? It explicitly maps its philosophy to the heaviest hitters in global regulation.

Host: It really does. The external validation here is very robust. The document points first to NIST, the National Institute of Standards and Technology in the United States.

Co-host: And what do they say?

Host: Well, NIST's AI Risk Management Framework explicitly advocates for context and risk-based AI governance. It categorically rejects a single, undifferentiated control approach.

Co-host: Which means NIST is effectively rejecting that binary block it or automate it light switch.

Host: Precisely. You cannot govern an enterprise neural network with just a blanket policy. You are required to assess the context and risk of each specific application, which is exactly the mechanical function of setting a justified target per capability on the AEOM dial.

Co-host: Right. The framework is essentially a direct operationalisation of NIST's guidelines. And furthermore, the text cites the official EU AI Act.

Host: And they are not messing around.

Co-host: No, they are not. The EU legislation mandates a legal framework covering governance, documentation, transparency, monitoring, and human oversight for relevant AI systems.

Host: Okay, let's look at the outputs of the five discovery questions we just analysed.

Co-host: Yeah, match them up.

Host: Requiring a deterministic receipt for every action, that perfectly satisfies the documentation and transparency mandate.

Co-host: Yes, it does.

Host: And defining the exact tripwires, revocation, that completely satisfies the monitoring and human oversight mandate.

Co-host: Exactly. The AEOM maturity dial totally transcends just operational efficiency. It is a core compliance mechanism.

Host: Global standards bodies are demanding that organisations govern AI based on segmented risk and maintain verifiable Human Sovereignty.

Co-host: And this framework gives you the exact operational blueprint to actually satisfy those incoming legal demands.

Host: It does.

Co-host: The implication there is honestly kind of severe.

Host: Very severe.

Co-host: If an enterprise is sitting there today, completely paralysed by this binary debate, relying on a blanket block it all policy while shadow IT scales in the background, they aren't merely sacrificing productivity.

Host: No, they're not.

Co-host: They are flying entirely blind into heavily regulated airspace. And without a maturity model that dials autonomy based on capability and risk, they will inevitably face catastrophic governance exposure when these international regulations take full effect.

Host: The failure to implement an operating model is no longer just a technical oversight.

Co-host: It is genuinely a failure of fiduciary duty at this point.

Host: So if we step back and view the macro picture here, we started this deep dive analysing a moment of pure operational panic.

Co-host: The rogue intern.

Host: Right. The machine executing complex actions autonomously and the enterprise completely freezing because it lacked the architecture to govern it.

Co-host: Yeah. And we examined how the absence of an operating model traps businesses in that binary debate, stifling safe automation while dangerously ignoring those shadow deployments.

Host: And we then unpacked the paradigm shift required to actually solve this, which is replacing the binary switch with the AEOM 1 to 5 maturity dial.

Co-host: The dimmer switch.

Host: The dimmer switch. Focusing on a justified target per capability anchored in that concept of Human Sovereignty, organisations can actually compartmentalise their risk.

Co-host: Yeah.

Host: We analysed the five discovery questions that force a business to rigidly define value, risk, systemic evidence, transactional transparency, and those all-important revocation tripwires.

Co-host: The circuit breakers.

Host: Exactly. And finally, we confirmed that this highly structured approach is directly aligned with the incoming legal realities of the EU AI Act and NIST standards.

Co-host: It's all there.

Host: And the document concludes with a highly practical call to action, which I think is great.

Co-host: What's the takeaway for the listener?

Host: It basically advises against attempting to map your entire enterprise architecture overnight.

Co-host: Don't boil the ocean.

Host: Right. The objective is to take just a single business capability, run it through the discovery questions, place it on the AEOM 1 to 5 maturity dial, and engineer its justified target.

Co-host: You pilot the framework, not just the technology.

Host: Yes. So to everyone listening, look at your own operational dashboards this week. Try to identify just one high volume, low judgment task that's causing friction in your department. Map that single task to this framework. Find the exact evidence the agent would need to produce to move up the dial, and find the specific circuit breaker that returns authority to you. If you build the operating model for just one capability, the path to safely scaling autonomous value actually becomes clear.

Co-host: Because it is the only methodical way to reclaim sovereignty over the system.

Host: It really is.

Co-host: But before we wrap up today, there is one final operational reality I really want you to consider. And it builds directly on that 5th discovery question, the revocation protocol.

Host: The tripwire.

Co-host: The tripwire, yes. We established that returning authority to a human decision maker is the ultimate safeguard of Human Sovereignty.

Host: Right, it's the circuit breaker that prevents total systemic failure.

Co-host: But consider the human psychology on the other side of that ledger for a second. Imagine you successfully implement this AEOM framework. You have an AI agent operating a complex supply chain capability at level 4 autonomy for, let's say, eight months.

Host: Okay, running smoothly.

Co-host: Yeah, it has produced flawless deterministic receipts. It's optimised thousands of shipments perfectly. You, the human operator, have completely tuned out of this workflow.

Host: Right, because it...

Co-host: It is entirely off your cognitive radar.

Host: Systemic reliability naturally breeds human complacency.

Co-host: It always does. Exactly. But then some unprecedented, bizarre geopolitical edge case occurs. The AI's predictive confidence score plummets, the circuit breaker trips exactly as engineered, and authority returns immediately to you, the human decision maker.

Host: Right.

Co-host: If you haven't actively analysed that specific workflow in eight months, are you cognitively prepared to catch the steering wheel at highway speeds?

Host: Probably not.

Co-host: How quickly can a human genuinely regain sovereignty over a complex, high-velocity operation they've been totally divorced from? I mean, we are spending millions engineering systems that can safely hand control back to a human.

Host: But the real question is?

Co-host: The real question, the next massive enterprise challenge, might be ensuring the human is actually awake enough to take it.

Host: Wow. That's a scary thought.

Co-host: Keep that in mind the next time you push the dial up to level 4.

Next step

Want to see what this looks like on your own BPM content? One conversation is enough to start.

Talk to Gareth