logo

The pilot's purgatory: when AI never makes it to production

August 20, 2026

Most AI pilot projects don't fail: they get stuck. They work, impress in the demo, receive praise, and then never go into production. They remain in a limbo where no one formally cancels them—canceling them would be admitting a mistake—and no one deploys them—deploying them would require solving problems the pilot project avoided by design.

This limbo has a cost that does not appear in any report: consumed budget, spent internal credibility and, above all, the erroneous conclusion that "AI does not work in our sector".

 

How many pilots are left behind?

The figures vary greatly depending on who is measuring and how, so it's best to look at several sources instead of relying on a headline.

He IBM's 2025 CEO Study, The study, which surveyed 2,000 CEOs from 33 countries, found that only 251% of AI initiatives had produced the expected return, and only 161% had been scaled to the organizational level. It is the methodologically most robust of the three sources.

S&P Global Market Intelligence It was published in March 2025 that 42% of companies had abandoned most of their AI initiatives, compared to 17% the previous year, and that on average 46% of proof-of-concept projects were discarded before reaching production.

And then there's the piece of information that circulated the most: the report MIT Project NANDA According to the 2025 report, approximately 95.1% of organizations did not obtain a measurable return from generative AI. It's important to be honest about this last point: it's a preliminary, non-peer-reviewed report based on 52 interviews and 153 surveys, and it has received serious methodological criticism regarding its measurement window and sample size. The overall direction aligns with other sources; the exact magnitude remains undetermined.

The important thing isn't which of the three figures is correct. It's that all three point to the same breaking point: it's not in building the pilot, it's in the journey between the pilot and production.

The six reasons why a pilot doesn't cross

Cause

What did the pilot avoid?

What does production require?

Data

A clean and selected set

Real, incomplete and contradictory data

Permits

A user with full access

Different roles, limited visibility, traceability

Exceptions

The happy path

The 15-30% of rare cases that were not designed

Integration

Copy and paste between screens

ERP, CRM and identity integration

Property

An innovation team

A business manager and a support team

Cost

Test volume

Unit cost at real scale, with peaks

The data. The pilot runs on a prepared extract. In production, the data arrives incomplete, duplicated, and inconsistent between systems. The model doesn't visibly fail: it chooses one of the contradictory versions without warning.

The permits. During testing, the system sees everything. In production, a salesperson can't see margins, a technician can't see health data, and an external consultant can't see anything that isn't theirs. If the permissions model wasn't in the design, it's not an adjustment: it's a complete overhaul.

The exceptions. The pilot demonstrates the successful path because that's what can be proven in twenty minutes. The real operation is, to a large extent, managing what doesn't fit. That percentage—between 15% and 30% of the volume in most of the processes we've analyzed—is where automation breaks down and where decisions must be made about what to scale to a human and how.

Integration. A demo might allow the user to copy information between windows. A production system, however, must read the operational context and execute actions within existing systems, respecting permissions and leaving a log.

The property. Pilots are usually born into innovation or technology. Production requires a business owner: someone whose annual goal depends on that process running more smoothly. Without that owner, there's no one to advocate for the budget for the next phase or to decide on exceptions.

The unit cost. In the pilot phase, the cost is irrelevant because the volume is small. At full scale, with peak demand and retries, the cost per case can render a project that was brilliant in the demo unfeasible. This calculation should be done beforehand, not afterward.

We have developed the complete journey between both phases in From AI PoC to Production

How to design a pilot that can actually cross

The solution isn't to make pilot projects bigger. It's to do them with the production constraints from day one, albeit on a reduced scale.

Define the business metric before you begin. Not "number of inquiries" or "active users": cycle time, error rate, cost per case, conversion rate, or margin. And it defines the threshold that determines whether to continue or stop. A pilot program without a stopping criterion is not an experiment; it's a disguised compromise.

Use real data from the beginning, even if there is little of it. A pilot with one hundred real cases teaches more than one with ten thousand clean cases, because the one hundred real cases include the problems you're going to have.

It includes an actual exception in the scope. Choose the most frequent rare case and resolve it within the pilot. If the design can't handle one exception, it won't handle thirty.

Appoint the business owner before the technical team. And that it be someone with responsibility for the outcome of the process, not for the technology.

Calculate the unit cost at actual volume. Extrapolate from day one. If the number doesn't work out on a larger scale, it's best to find out by week two.

Put an expiration date on it. Four weeks is enough time to determine if something works, but too short for it to become a zombie project. That's precisely why our proofs of concept take this format: a fully functional PoC in one month, complete with guardrails, traceability, and human oversight. Not because it's commercially appealing, but because a short timeframe forces us to limit the scope to something that can truly be validated.

What to do with a pilot who has been stuck for months

If the pilot program already exists and is not progressing, there are three honest solutions, and none of them involve waiting.

Diagnose and decide. An analysis of model, data, integration, and cost that answers three questions: what can be salvaged, what needs to be rewritten, and what needs to be stopped. This is the objective of a AI project audit , and is usually resolved within weeks.

Reduce the scope to something that can actually be deployed. Often, the 20% of the original scope yields the 80% of value and is deployable within a month. The resistance to doing so is political, not technical: reducing the scope seems to admit partial failure.

Stop it and document why. It's the most undervalued option. A pilot who's stopped working with a clear record of what they've learned is an asset. A pilot in limbo is a liability that consumes attention every quarter.

The problem wasn't AI.

When analyzing stuck pilots, almost none died because of the model itself. They died because of the data, the permissions, the exceptions, the integration, or the lack of an owner. In other words, they died for the same reasons software projects have been failing for the last thirty years.

That's essentially good news. It means the problem isn't a technological mystery, but rather an engineering and project governance issue. And we do know how to solve those.

The question that should be taken to the next committee meeting is not whether AI works. It's what needs to be fixed in our data, our permissions, and our integration so that any automation, with or without AI, can reach production.

Frequently Asked Questions

Why do most AI pilots fail to reach production?

For six recurring reasons: incomplete and contradictory real data, a permit model that the pilot did not consider, exceptions not designed, lack of integration with existing systems, absence of a business manager and a unit cost that is not sustainable on a real scale.

According to IBM's 2025 CEO Study, only 251% of AI initiatives have delivered the expected return, and only 161% have been scaled. S&P Global estimated in March 2025 that 421% of companies had abandoned most of their AI initiatives. These figures vary considerably depending on the study's methodology.

Around four weeks. That's enough time to validate whether the case works with real data, but too short for the pilot to become an indefinite project with no end in sight. A short timeframe forces us to limit the scope to something verifiable.

Business metrics defined before starting: cycle time, error rate, cost per case, conversion rate, or margin, with an explicit threshold that determines whether the project continues or stops. The number of inquiries or active users does not prove business impact.

There are three ways out: diagnose the problem through an audit that determines what to salvage, what to rewrite, and what to stop; reduce the scope to the portion that can be deployed in weeks; or stop the project while documenting the lessons learned. Leaving it in limbo is the only option that generates costs.

That figure comes from a preliminary 2025 MIT Project NANDA report, which is not peer-reviewed and is based on a small sample, and has received methodological criticism. Other, more reliable sources, such as IBM, point in the same direction with different magnitudes: the scaling problem is real, but the exact figure is not yet established.

 

Has your AI pilot been sitting idle for months? We analyze your model, data, integration, and cost, and tell you what to save, what to rewrite, and what to stop. If the case is viable, we'll develop a working Proof of Concept (PoC) in four weeks, complete with safeguards and traceability. Let's talk about your case →

Artificial Intelligence project stalled in pilot phase before reaching production