ai
Why Most AI Pilots Fail
The demo works. The pilot dies. This pattern is so common it deserves a name, and if you have sat through enough of these you can feel the death coming from the second week: the slow slide from “this is amazing” to “let’s revisit next quarter” to silence.
Here is the part that surprises people. When you do the post-mortem on a failed AI pilot, you almost never find a model problem. The model was fine. The model did, in the demo, exactly the impressive thing everyone saw it do. What killed the pilot was one of two things, and usually both. It was the plumbing, and it was the goal.
The plumbing problem
A demo runs on a clean, curated example that someone prepared in advance. A pilot has to run on your actual data, and your actual data is a mess: scattered across a dozen systems, stale in half of them, and locked behind three approvals in the rest. The model was never the bottleneck. The bottleneck was that it could never get a clean, timely feed of the information it needed to do the job for real.
This is the uncomfortable truth under most “AI readiness” conversations. What people call AI readiness is really data readiness, and most organizations are not data-ready. The information the model needs lives in someone’s inbox, a spreadsheet on a shared drive that gets updated when somebody remembers, a SaaS tool whose export requires a ticket, and a database only one team can touch. In the demo, a human quietly gathered all of that and laid it out nicely. In production there is no human doing that every time, and the moment the model has to reach for the real data through the real plumbing, it starves.
You can have the best model in the world and it will produce nothing useful if the pipe feeding it is clogged, dry, or delivering last month’s numbers. The unglamorous work of connecting the systems, cleaning the feed, and making the right data reachable automatically is not the boring prelude to the AI project. It very often is the AI project. The model was the easy part.
The goal problem
The second killer is quieter and even more common: nobody agreed on what “working” actually meant before the pilot started.
A demo answers an easy question. Can the model do this thing? A pilot has to answer a much harder one. Can our organization do this, repeatedly, with our data, inside our real processes, well enough to matter? Those are entirely different bars, and teams that never define the second one drift for a couple of months and then quietly conclude the pilot “didn’t really pan out,” because there was never a line that would have told them whether it did.
So before you start the next one, force answers to three questions, and treat a bad answer to any of them as a reason to stop before you spend the money:
- What decision or task does this replace, exactly? Name the specific thing a specific person does today that this will do instead. If the honest answer is “it helps generally” or “it makes us more efficient,” stop. Vague scope is the single most common cause of a pilot that cannot be evaluated, because you cannot tell whether a fog got any thinner.
- Where does the input data live today, and who owns it? Trace the actual data the thing needs back to its home and its owner. If the answer involves a quarterly export, a person who “usually pulls that,” or a system nobody controls, you have found your plumbing problem before it kills you. Fix the pipe first, or pick a use case whose data is actually reachable.
- What number moves if it works? Name the metric, whether it is time saved, error rate, throughput, cost, or revenue, that will visibly change. If no number moves, that is fine, but be honest about what you have. It is a toy. And toys are genuinely fine. Playing with the technology to learn it is a legitimate use of a little money. Just budget them as toys, not as bets, and do not act shocked when a toy does not transform the business.
The boring work is the work
Here is the pattern I have watched hold across every one of these. The teams that get pilots into production are rarely the ones with the best models or the flashiest demos. They are the ones who did the boring integration work first: who connected the systems, cleaned the data feed, named the metric, and scoped the task down to something specific and real before they got anywhere near the exciting part.
There is a companion failure that is about people and adoption, and I have written about that one separately, the organizational immune system that quietly rejects a tool nobody was brought along on. But even when the people are perfectly willing, the plumbing and the goal will sink you on their own, silently, without anyone resisting anything. The two failures stack. A pilot with clean data and a clear metric can still die if the humans feel threatened, and a pilot with fully committed humans still dies if the data never arrives. You have to clear both.
Start smaller than feels impressive
One more thing, because it follows directly from all of the above. The instinct in a lot of organizations is to make the first pilot ambitious, so it will be worth the effort and look good to leadership. That instinct is backwards. An ambitious first pilot maximizes exactly the two risks that kill pilots: it needs data from more systems, so the plumbing problem is worse, and it tries to do something fuzzy and grand, so the goal problem is worse.
Pick something almost embarrassingly small and specific instead. One task, one clear metric, data that is already reachable. Prove that your organization can actually get an AI capability into daily use and keep it there. That first small win teaches you where your plumbing leaks and how your people react, and it earns you the credibility to attempt something bigger. Chains of small, real wins are how the ambitious version eventually gets built.
What “production” actually means
It helps to be clear-eyed about what it takes for a pilot to become a system, because “it worked in the pilot” is not the finish line either. Production means someone owns it after the excitement fades and the budget line goes quiet. It means monitoring, so you notice when the data feed silently breaks or the outputs slowly drift away from useful. It means a human fallback for the cases the model gets wrong, because it will get some wrong, and a system with no plan for its own errors is a liability wearing a demo’s clothes. And it means the process around the tool actually changed to use it, rather than the tool sitting off to the side while everyone keeps doing the work the old way out of habit.
Most pilots that technically succeed still never reach production, because nobody planned for the boring, permanent work of keeping the thing alive once it stopped being a shiny project. So ask, before you start: who runs this in a year, when it is just part of how the work gets done and no longer has a champion or a launch date? If the honest answer is nobody, you are not building a system. You are building a demo with a longer runway, and it will die the same quiet death as all the others, just later.
So do the unglamorous work up front. Define exactly what the model replaces, make sure it can actually reach the data it needs, and name the number that will move. Do that and your pilot has a real chance of becoming a system. Skip it, and you will have another great demo, another dead pilot, and another post-mortem that, once again, will not find a single thing wrong with the model.