Why 95% of AI pilots fail, and how FSI Frontier Firms turn AI into business value
Ask any CDIO in banking or insurance about their AI roadmap and you’ll probably get some real enthusiasm. Ask about pilot number four – or the one that stopped showing up in programme updates at around month three – and the tone changes. MIT puts the failure rate for AI pilots at 95%. It’s a number that is quoted a lot, often as a warning. Instead, it could be treated as something more useful: a description of what’s happening across financial services and insurance – rather than a hypothetical risk to plan around.
Maybe you’re familiar with the pattern. A pilot starts with impressive momentum, has its own slide in a company update. Then the questions start to trickle in, the data turns out to be siloed after all, and the project drifts into “still evaluating” purgatory where it lives until someone finally archives the teams channel. We’re not being purposefully negative, and this isn’t a failure of ambition. The CDIOs we talk to know exactly what they want AI to do. The gap they experience is operational. Organisations are often trying to introduce AI into legacy processes that were never designed for it, while governance, security and workforce adoption struggles to keep pace. The result is disconnected pilots, isolated automations, and technology that never becomes part of how work actually gets done.
It looks different, depending on where you sit
Banking’s version of this is fraud, and AML pilots that stall the moment governance requirements crop up. Insurance is trying to modernise claims and underwriting while also fixing how it talks to policyholders. Capital markets is exploring AI in trading and risk modelling, but the data sensitivity means every security question is asked twice, sometimes more. Different sub-sector, same shape of problem: good intentions, real platforms, and a structural reason why none of it adds up to anything yet.
What the 5% are doing instead
A handful of organisations have managed to move beyond pilots and into scaled adoption. The pattern is consistent enough to be described.
They start with a problem, not a tool
The best, and most successful, initiatives didn’t begin with “let’s try AI somewhere”. They began with a business outcome or operational pain point, then assessed where AI, automation and people could create the greatest impact, balancing value against feasibility, risk, and regulatory fit. Curiosity projects rarely survive that filter. The strongest programmes start with a top-down view of customer outcomes, operational priorities and growth objectives before assessing the current state.
They put AI inside the workflow, not bolt it onto the homepage
A chatbot sitting on a website doesn’t change how a business runs. The organisations seeing traction are embedding AI into the tools people already use every day: Dynamics 365 surfacing insights and recommended actions mid-conversation, Copilot and Cowork supporting knowledge works in daily tasks, and Power Platform automating claims triage. The work moves becomes part of daily operations instead of sitting beside them, making adoption far more likely.
They use what they’ve already paid for
This is the one that tends to surprise people. Most teams don’t need to build bespoke AI from the ground up; they need to unlock the AI capability already sitting inside Microsoft 365, Dynamics 365, Power Platform, Azure, and rethink how those platforms can support AI-enabled ways of working. One Head of Business Applications at a major financial services firm put it plainly: they were surprised by how much they could already do with systems they already owned.
They let the people closest to the problem experiment
The best use cases often come from the relationship manager, the claims handler, the compliance officer – people who see where a process slows down because they are the ones stuck in it. Frontier firms give these teams ringfenced time to test ideas, supported by training, enablement and clear governance that allows experimentation without creating new risks.
They budget like they mean to scale
The gap between a pilot and a platform is usually a planning gap. Too many initiatives get funded for a single proof of concept, with nothing set aside for the security, training, and integration work that production needs. The teams that scale plan for that from day one, including the investment needed for security, Responsible AI, integration and workforce adoption.
What it all adds up to
One UK bank had more than 40 Power Apps live with no central governance, which is innovation in name only; nobody could see what existed or stop it duplicating. Elsewhere, a finance team automated a reporting process through Power Platform and got back more than 200 hours a month they’d been spending on reconciliation instead of analysis. Same starting ingredients – ambition, platforms, and people – different outcomes, because the second team treated AI as an operating model transformation rather than a collection of experiments. If ‘let’s look at pilot four’ is sounding familiar, the issue probably isn’t your AI strategy. It may be that you’re applying AI to processes, platforms and ways of working that were designed for a different era. The organisations creating value are redesigning how work gets done, with AI, Copilot, automation and increasingly agents built into the operating model from the start.
We’ve written a full breakdown of what’s blocking AI adoption across banking, insurance, and capital markets, including how FBD insurance worked through it in our eBook: The AI Disconnect in Financial Services and Insurance. Download your copy here.
