logo

The ROI of AI is not measured in prompts

October 6, 2026

The number of queries, active users, or executed prompts doesn't measure the return on investment in artificial intelligence: it measures its use. These are different things, and confusing them is why so many projects are presented as successful to the committee while the bottom line shows nothing.

A system with high adoption and zero return is perfectly possible. In fact, it's the most common scenario.

Why usage metrics are misleading

A usage metric answers the question, "Are people using it?" A return on investment metric answers, "Is the company making a profit?" There's a gap between the two that's almost never documented.

The typical case: a tool saves eighty people fifteen minutes a day. On paper, that's twenty workdays a month. In practice, it's fifteen minutes spread out and absorbed into other tasks, without any additional budget line item appearing.

McKinsey quantified the phenomenon in its 2025 survey of 1,993 participants in 105 countries: 88.1% of organizations use AI in some function, but only 39.1% report an impact on EBIT, and barely 6.1% exceed 5.1% of attributable EBIT. High usage, rare return.

The three ways in which saving does lead to results

Any serious business case has to explain which of these three it is pursuing:

A whole step disappears. A task isn't accelerated: it ceases to exist. A validation that is no longer needed because the data is checked at the source.

More volume is absorbed without expanding the structure. The team manages one more 30% case with the same staff. This is the most common scenario in growing companies and the easiest to defend.

An identifiable direct cost is reduced. Fewer errors, less rework, fewer contractual penalties, less overtime at month-end closing.

If a project cannot specify which of the three it is seeking, it does not have a business case. It has a reasonable expectation, which is something else entirely.

The four metrics that a committee should require

Metrics

What does it measure?

Why it matters

Cycle time

From the moment a case is entered until it is closed

It reflects the entire process, not just the task.

Cost per case resolved

Total cost divided by closed cases

It is comparable to the previous process

Error or rework rate

Cases that need to be corrected afterwards

Detects savings that are paid elsewhere

Absorbed capacity

Volume managed with the same structure

Translate savings into growth

The third option avoids the most expensive trap. A system that resolves cases for €0.40 compared to €3 for the manual process seems like a resounding success, until it's discovered that the 20% requires a subsequent correction of €5. The net result is different, and only becomes apparent if someone measures it.

The baseline: what to do before you start

None of these metrics are useful without the previous value. And the time to capture that value is before making any changes, because afterward, it can't be reconstructed without discussion.

Four data points, measured over two or three weeks:

  1. How many cases are processed during the period.
  2. How long does each one take?, including waiting and coordination, not just active work.
  3. How many require correction later and what that correction costs.
  4. How much does the entire process cost? For example, with the imputed personnel cost.

This job has a valuable property: It produces results even if the project is not carried out.. Measuring a process reveals bottlenecks that are sometimes corrected without any technology.

The decision threshold, written before

The second discipline that separates successful projects from those that drag on forever is to set in advance the number that determines whether to continue or stop.

It must meet three conditions: be a business metric, have a specific value, and have a deadline. "If the cost per case doesn't fall below X in eight weeks, we'll stop" is a threshold. "We will continuously evaluate the impact" is not.

Without a written threshold, the evaluation only happens when too much has already been invested to admit that it's not working. This is the mechanism that produces the permanent pilot programs we described in the pilot's purgatory

The cost that almost no one includes

An honest return calculation should subtract five items, not one:

  • He consumption of the model, including retries and peaks.
  • He supervision time, which is reduced but does not disappear.
  • The continuous assessment, to find out if the system is still working.
  • The operation and maintenance, which exist even if nothing fails.
  • The bug correction and its impact on the customer.

Presenting a return that only subtracts consumption produces a number that won't survive the first serious review. We've developed this in Budget for AI as a product, not as a campaign

The one-line question

In any AI project results report, there's one question that clarifies the conversation in seconds: Which budget line has changed?

If there's a response—this item has decreased, this team is handling more volume, these penalties have disappeared—there's a return. If the response describes user satisfaction, the number of inquiries, or a perceived agility, there's adoption.

Adoption is a prerequisite for return. It is not the return.

Frequently Asked Questions

How do you measure the ROI of an artificial intelligence project?

With four business metrics: complete process cycle time, cost per resolved case, error or rework rate, and capacity absorbed with the same structure. The number of users, queries, or prompts measures adoption, not return.

Because fifteen minutes saved for eighty people are absorbed into other tasks without generating any budget line. Savings only become a result if a step in the process is eliminated, if more volume is absorbed without expanding the structure, or if an identifiable direct cost is reduced.

This is the state of the process before intervention: number of cases, time per case including waiting periods, percentage requiring correction, and total cost per case. This data must be captured for two or three weeks before any changes are made, because it cannot be reconstructed afterward without discussion.

Five: model consumption with retries and peaks, human supervision time that is reduced but not eliminated, continuous performance evaluation, operation and maintenance, and error correction with its impact on the customer.

It's the value that determines in advance whether a project continues or stops. It must be a business metric, have a specific number and a date: for example, if the cost per case doesn't fall below a certain figure within eight weeks, the project is stopped.

Which budget line item has changed? If the answer identifies a line item that has decreased, a team that is handling more workload, or penalties that have disappeared, there's a return on investment. If it describes user satisfaction or the number of inquiries, there's adoption.

Does your AI project have a baseline and decision threshold? We define the metrics, measure the current state, and set the number that will determine whether it continues. Two hours of analysis, no commitment. Let's talk →

Artificial Intelligence provider analyzing the real needs and processes of a company
Company strengthening its digital sovereignty through Artificial Intelligence, open technological architecture and integration of business platforms.
Business leadership in enterprise Artificial Intelligence projects