IGX Solutions
Podcast

Exposing Invisible Single Points of Failure

Local continuity plans can all look complete while a single supplier, system or person that several critical services quietly depend on stays invisible at enterprise level. This episode covers why that dependency picture only surfaces after a disruption, what it costs to reconstruct it mid-incident, and what a connected, queryable dependency model replaces it with.

Episode 6 IGX360

Episodes feature AI-generated hosts discussing human-written IGX360 research.

In this episode

Local continuity plans can each look complete on their own while a shared supplier, system or individual that several critical services quietly depend on stays invisible at enterprise level. This episode covers why that dependency picture only becomes visible once a disruption crosses shared technology, suppliers, locations, workforce or service boundaries faster than anyone can trace the impact, and why that gap turns the first real incident into the first real dependency test.

Drawing on FCA operational resilience requirements to map important business services and stay within impact tolerances, and ISO 22301’s framework for identifying disruption risk and maintaining continuity, the conversation walks through why concentration and single points of failure need to be visible before disruption, not reconstructed during it, and how a connected process model with automated dependency mapping turns scenario testing and resilience investment into something targeted rather than guesswork.

Read the full transcript

Host: So imagine a multi-billion dollar enterprise.

Co-host: Okay, I'm picturing it.

Host: And one day it just goes completely dark. But here's the crazy part: the marketing department's backup plan, it worked perfectly. Human resources, their contingency plan deployed without a hitch. IT ticked every single box on their localised business continuity checklist. I mean, everyone did their job flawlessly.

Co-host: But the entire company is still in absolute free fall.

Host: Exactly.

Co-host: Completely paralysed.

Host: Yeah, I mean, it sounds impossible on paper. You just assume that, well, if every individual component has a safety net, the whole structure has to be secure. It makes logical sense.

Co-host: It does, but when you look at modern operational architecture, that assumption is not just flawed. It is actually the exact reason massive systems collapse.

Host: And that is exactly what we are getting into today. Welcome to today's deep dive.

Co-host: You, our listener, have shared a remarkably intriguing corporate strategy document with us. It is a fascinating read.

Host: It really is. It's titled simply P5: We do not know where operational failure will spread.

Co-host: Very ominous title.

Host: Right, very to the point.

Co-host: So our mission today is to unpack this roadmap, understand the anatomy of these unpredictable failures, and explore how a company can actually expose its hidden vulnerabilities before they trigger that total collapse.

Host: And to set the stage for you, the source material tackles this using the SPIN framework: situation, problem, implication, need payoff.

Co-host: I love a good framework.

Host: Oh, absolutely. And it ends with specific discovery questions and real-world validation.

Co-host: It's essentially a very structured roadmap for exposing a company's blind spots. And it highlights a core paradox in business operations today.

Host: Which is?

Co-host: Well, we operate in an environment where our dependency is on suppliers, internal systems, specialised roles, automated processes, and they are all entirely fragmented.

Host: It's totally disconnected.

Co-host: Exactly. The data mapping these connections is either incomplete or, honestly more dangerously, it's just too difficult to query in real time.

Host: Meaning you can't just type into a search bar: what happens to the company if this one specific vendor goes offline?

Co-host: Right, because the system itself doesn't actually know all the tentacles attached to that vendor. Those dependencies cross multiple boundaries.

Host: The text points out that when a disruption occurs, it doesn't respect departmental lines. It doesn't care if you're in HR or marketing.

Co-host: Exactly. A failure jumps across shared technology stacks, it crosses into different supplier tiers, and it ripples through geographical locations and workforce boundaries far faster than the actual impact can be tracked by the people watching it unfold.

Host: Okay, let's ground this in a visual. Earlier I was thinking about this like a tangled ball of holiday lights.

Co-host: Oh, that's a good one.

Host: Right, like if one bulb goes out, you can't just trace the wire, because the disruption crosses over ten different strands before you even realise half the tree is dark.

Co-host: Yeah, but even that might not capture the structural danger we're really talking about here.

Host: Really? How so?

Co-host: Well, the problem identified in the text is that these single points of failure remain deeply embedded in the operating model, because there is an absolute absence of an end-to-end view.

Host: No one sees the big picture.

Co-host: Right, no one can see the full picture of critical outcomes, concentration, substitution, controls, recovery assumptions.

Host: So it's more like, okay, think of a massive newly constructed skyscraper.

Co-host: Okay, tracking with you.

Host: Every single floor of that building, which represents a different department, hires its own interior designer, and they all install top-of-the-line fire extinguishers on their floor.

Co-host: Right, so from the perspective of each floor, they are perfectly safe.

Host: Yes, the local departments are just looking at the cosmetic paint and the local fire alarms. They have no idea that the plumbing for their floor, the electricity for the floor above them, and the server cooling system for the basement all run through the exact same structural column.

Co-host: So if that one column buckles, the whole building comes down and all those local fire extinguishers are completely useless.

Host: That is a perfect analogy. That structural column is the embedded single point of failure.

Co-host: The underlying weakness isn't that companies don't try to plan for disasters.

Host: Right, you spend millions on it.

Co-host: They really do. But the problem is that absolute absence of the end-to-end view. Which brings us to a quote in the source material that legitimately stopped me in my tracks.

Host: I think I know which one you're going to say. It's so jarring.

Co-host: The document states, quote: an incident becomes the first full dependency test.

Host: It is a really sobering realisation.

Co-host: It is terrifying. I mean, think about what that actually means mechanically. You are jumping out of an aeroplane, but you have no idea if your backup parachute is packed correctly until you pull the cord in free fall.

Host: Yeah, the disaster itself is your very first real-world stress test. That is just wild to me.

Co-host: It is, and it explains why recovery from these outages is always so much slower than executives expect.

Host: Because they're figuring it out on the fly.

Co-host: Exactly. During a major incident, the impact on customers and the violations of regulatory compliance just continue to snowball. The teams are trying to map the dependencies in real time.

Host: They're trying to read the blueprints of the building while the building is actively on fire.

Co-host: Yes, exactly that.

Host: I have to push back here though. I want to play devil's advocate for a second on behalf of our listeners.

Co-host: Sure, go ahead.

Host: We are talking about massive enterprises here. Banks, global supply chains, tech giants. They have entire floors dedicated to risk management, they employ thousands of auditors. So how can an incident be their first full test? Surely these multi-billion dollar companies know where their own structural columns are, right?

Co-host: You would definitely think so, but the document addresses this directly, and the answer comes back to those so-called complete local plans.

Host: The fire extinguishers on every floor.

Co-host: Exactly. Department-scale localised continuity plans give the illusion of total security.

Host: So the auditors come in, they look at the marketing department's thick binder of protocols, and they sign off. Check. And they look at HR's protocols, and they sign off.

Co-host: But they are auditing the floors, not the whole building.

Host: Right.

Co-host: Let's look at the mechanics of how this happens in practice. Say the marketing department's backup plan relies on routing their emergency communications through a specific, highly secure cloud service provider. Meanwhile, three floors down, HR's backup plan involves routing their emergency payroll processing through a totally different third-party vendor.

Host: Okay, let me guess. That third-party payroll vendor also relies on the exact same cloud service provider that marketing is using.

Co-host: Bingo. That is the invisible structural column. Neither department knows about the other's underlying infrastructure.

Host: Wow. Locally, both look completely protected to any auditor.

Co-host: But at the enterprise level, there is a massive unseen concentration of risk. Because if that one cloud provider goes down, marketing and HR both fail simultaneously. Their shared single point of failure remained completely invisible precisely because everyone was so confident in their siloed local plans.

Host: So the illusion of safety actually manufactures the vulnerability. If the local plans are essentially lying to the company about its overall resilience, how do we fix this without tearing the whole company down and starting over?

Co-host: Well, the document proposes a specific operational shift: building a connected process model. The objective of this model is to reveal that concentration of risk, the dependency overlaps, and the control exposures, well before a failure ever actually happens.

Host: But wait, aren't we just talking about creating another massive, bloated IT auditing project? Because every few years a company decides to map its processes and it turns into this two-year consultancy nightmare, and by the time they finish mapping the data, the technology has changed and the map is entirely obsolete.

Co-host: Oh, I've seen that happen so many times. Not exactly, because a connected process model isn't a static drawing or a one-time audit. It is dynamic.

Host: Like a living document.

Co-host: More like a live digital twin of your operation. So instead of mapping every single process in the company equally, it focuses specifically on the dependencies that have the potential to exceed what the document calls impact tolerances, or break service commitments.

Host: Let's pause and define those two terms, because they seem really critical here. When the document mentions impact tolerances, what are we actually measuring?

Co-host: Impact tolerance is the maximum acceptable level of disruption to a crucial business service before the consequences become intolerable.

Host: Intolerable meaning what exactly?

Co-host: Think of it as the absolute threshold where a bad day turns into an existential crisis. For a bank, the impact tolerance for its mobile app being down might be, say, four hours.

Host: Right, people need to see their money.

Co-host: Exactly. Beyond four hours, the financial loss and the reputational damage just become catastrophic.

Host: And what about service commitments?

Co-host: Those are your service level agreements, or SLAs. It is what you are contractually obligated to provide to your customers or partners. So if your contract says payment processing will never be down for more than sixty minutes, that is your service commitment.

Host: Exactly right. Okay, so tying it all together, what does this all mean for the listener?

Co-host: It's kind of like moving from studying a bunch of isolated flat 2D maps to looking at a live 3D hologram of a city's traffic grid, where you can finally see how all the intersections actually connect.

Host: That is a great way to picture it.

Co-host: And that targeted approach completely changes how a business allocates its financial investments.

Host: Also?

Co-host: Well, you stop spreading your resilience budget evenly across every department, like peanut butter, and instead you start focusing entirely on fortifying those major intersections.

Host: The structural columns.

Co-host: Exactly. You identify the single point of failure before the server actually catches fire.

Host: You are doing preventative structural reinforcement rather than just buying more fire extinguishers for people who don't need them.

Co-host: Yes. And it leads to much faster disruption impact analysis.

Host: Right, because you already have the map.

Co-host: Because if a vendor does go down, you aren't guessing who relies on them. The connected model instantly highlights every single downstream process that will be affected.

Host: Which is huge.

Co-host: It is. It also provides highly credible evidence for testing and recovery strategies, which is absolutely vital when you have to look your stakeholders in the eye and prove you are actually prepared.

Host: And the document actually provides a mechanism to test that preparedness, right?

Co-host: It lists several discovery questions meant to act as a diagnostic tool, and they are tough questions.

Host: So I want to turn this directly to you, our listener. Apply these questions to your own workplace, your own complex project, or your own supply chain. The first question the text poses is: which processes depend on just one supplier, system, or individual?

Co-host: I want to focus on the last word, individual. The individual dependency is notoriously difficult to spot.

Host: Oh, because everyone focuses on the tech.

Co-host: Right, companies focus heavily on technology and completely ignore the fact that an entire critical workflow might rely entirely on one senior engineer who just holds all the institutional knowledge in their head.

Host: The classic bus factor.

Co-host: If that person wins the lottery tomorrow and quits, does your service commitment fail?

Host: Wow, yeah. Okay, question two: can you identify downstream impacts before approving a change? This one is huge for IT. I know from experience how incredibly rare it is to actually have an answer to this.

Co-host: A team will push a software update, completely confident in their local environment, and then just wait to see if anyone three departments over starts screaming that their workflow is broken.

Host: Which really highlights a fundamental tension between agility and stability.

Co-host: Right, move fast and break things.

Host: Exactly. Companies want to move fast and push updates rapidly. But if you cannot see the downstream impact of a change before you make it, your agility is functionally indistinguishable from recklessness.

Co-host: That is a great way to put it. Okay, question three from the document: when was the dependency model last tested?

Host: And we are talking about testing the local backup server here.

Co-host: No, definitely not. We are talking about testing the actual map of how everything interconnects. If a company cannot produce a recent date for that end-to-end test, then we are right back to the terrifying reality we discussed earlier.

Host: The incident becomes the first test.

Co-host: The next incident will serve as their first full dependency test. Then we hit question four, which honestly I think is the most revealing metric in the entire framework: which dependency is shared by the greatest number of critical services?

Host: That question cuts straight through the illusion of local planning.

Co-host: Because nobody knows. If you gather five different department heads in a room and ask them that question, none of them will have the answer.

Host: They only know their own lane, their own vertical.

Co-host: But once you implement a connected process model, you can instantly see that, oh wow, five distinct critical services are all secretly funnelling through the exact same external database.

Host: And that database is your ultimate single point of failure, just hiding in plain sight.

Co-host: Exactly. Now, if implementing this level of rigorous mapping sounds like a nice-to-have theoretical exercise, the document brings in some heavy external validation to prove otherwise.

Host: Yeah, the regulators are no longer treating this as optional.

Co-host: No, they are not. The text points directly to the FCA, the Financial Conduct Authority in the UK, which sets a massive global precedent. They strictly require in-scope firms to physically map out their important business services.

Host: And they don't just want a nice flow chart, do they?

Co-host: No, definitely not. The FCA requires these firms to prove, with actual data, that they can remain within their impact tolerances during a severe but plausible disruption.

Host: So you cannot just tell a regulator, hey, don't worry, we have a backup plan.

Co-host: Not anymore. You have to demonstrate the mathematical reality that your critical services won't collapse if a tier-3 supplier suddenly files for bankruptcy. The document also anchors this in ISO 22301.

Host: For anyone listening who isn't deep in the compliance trenches, ISO 22301 is the global gold standard for business continuity management systems.

Co-host: The heavy hitter. It is a definitive framework for identifying risks, planning the response, and ensuring you can maintain your priority activities regardless of what chaos is happening in the background.

Host: And both the FCA and ISO 22301 demand exactly what this document advocates: eliminating siloed planning.

Co-host: Exactly. Adopting a connected, end-to-end understanding of your vulnerabilities. The challenge really remains: how do you actually execute this? I mean, if you try to map every single process, software, and vendor across an entire global enterprise all at once, your teams will be completely paralysed by the sheer volume of data.

Host: Oh, absolutely, it's overwhelming.

Co-host: Which is why the call to action at the end of this document is so sharply focused. It doesn't tell you to map the whole company. It says: book a call to expose the single points of failure around one critical business service.

Host: Limiting the scope is really the only way to begin.

Co-host: If you select just one vital service, for example processing customer payments, you can trace that specific operational strand from start to finish. The document mentions a specific technological route for doing this, combining IGX360 Insights with iGrafx.

Host: For clarity, let's explain what those actually are, because they aren't just corporate buzzwords.

Co-host: Right, so iGrafx is a business process management platform. It allows a company to build that digital twin we were talking about earlier.

Host: The live 3D map.

Co-host: Exactly, the live digital mapping of how work actually flows through an organisation. It visualises the processes. And then IGX360 Insights acts as the analytical overlay on top of that.

Host: Wow.

Co-host: Precisely. It applies resilience and risk lenses directly onto that digital map. You aren't just looking at a diagram of how payment processing works, you are looking at a live model that highlights exactly where the risk is concentrated.

Host: So by applying these tools to just one critical service, you uncover the hidden suppliers, the shared servers, and the single individuals that the payment process secretly relies on.

Co-host: You essentially light up one clear, navigable pathway through the murky waters. And once you prove the value on that single service, once you expose a massive blind spot the company didn't know it had, you now have a defensible blueprint to scale that mapping to the rest of the enterprise.

Host: You aren't boiling the ocean, you are securing the most important trade routes first.

Co-host: I love that. It transforms a theoretical risk exercise into an actionable, measurable operational upgrade.

Host: It really does. This entire document forces a fundamental shift in how we think about preparedness.

Co-host: But as we wrap up this deep dive, we want to leave you with something to really chew on. There is one final discovery question embedded in the source material that we haven't thoroughly dissected yet, and it might be the most profound challenge in the entire text.

Host: I think so too. The question is simply this: which recovery assumption has not been tested end to end?

Co-host: A recovery assumption, let's unpack that. Well, we spend massive amounts of capital building complex operating models. We invest in redundant servers, alternate supply chains, specialised continuity teams.

Host: All the fire extinguishers.

Co-host: Right. But in doing so, we inherently develop recovery assumptions. These are the things we treat as facts simply because they exist on a spreadsheet, or because we pay a premium for them. Like we assume our secondary cloud provider actually has the bandwidth to handle all of our traffic if the primary goes down.

Host: Exactly. We assume our external vendor will prioritise our support tickets during a massive regional blackout. We assume the backup generator has been maintained properly.

Co-host: Right. But if an incident truly is the first full dependency test, we have to face a very uncomfortable possibility.

Host: Which is?

Co-host: Is it possible that our greatest operational vulnerability isn't a lack of planning at all? Is it possible our greatest vulnerability is simply our blind faith in untested assumptions?

Host: That is a staggering perspective. We all want the clean, binary X-ray that points exactly to the broken bone.

Co-host: But we are operating in a world where the X-ray machine itself might be drawing power from an untested grid, relying on a software patch from a vendor that went out of business yesterday. The hidden dependencies are everywhere, and the assumptions we make about them are the true single points of failure.

Host: So as you step back into your own work today, take a hard look at the projects you manage and the systems you rely on. What is the one recovery assumption you are treating as an absolute certainty, solely because it hasn't been forced to fail yet?

Co-host: It is a question that should genuinely keep every operational leader up at night.

Host: Absolutely. Well, thank you so much for bringing such an incredibly thought-provoking and dense document to the table.

Co-host: We love unpacking these structural puzzles with you. Until next time, keep asking the hard questions, and whatever you do, do not trust the map until you have actually walked the territory.

Next step

Want to see what this looks like on your own BPM content? One conversation is enough to start.

Talk to Gareth