Article

How to measure AI ROI: From AI adoption to business impact

Last updated 
Sep 10, 2026
8
 min read
Episode 
8
 min
Published 
Sep 10, 2026
8
 min read
Published 
Sep 10, 2026
8
min

Picture a mid-market insurance company that rolled out an AI tool to help adjusters triage incoming claims. The rollout went smoothly. Adjusters logged in, used the tool, and kept using it. Six months later, leadership has a dashboard full of usage numbers and a comfortable answer to "are we doing AI?"

Ask the COO whether that investment is actually paying for itself, and the room gets quieter.

 Chart illustrating the adoption rates of ROI, displaying numerical data trends over time.

We see this measurement gap across most of the mid-market companies we work with. Conversations about AI return on investment (AI ROI) tend to center on a single question: is the tool being used? This question is easy to answer and rarely the one that matters. The harder question, the one that actually determines whether the investment made sense, is whether what you put into it produced a business outcome worth more than it cost.

A meaningful view of AI ROI has to connect four things: what you invested, what changed, what business outcome came out of that change, and what it costs to keep that outcome going. Miss any one of those, and the number you're reporting to your board is really just an educated guess with a spreadsheet attached.

Why most AI ROI numbers are guesses

The scale of this problem is well documented at this point. Eighty eight percent of AI pilots never reach production¹, and among the ones that do ship, MIT's 2025 review of enterprise deployments found that 95% show no measurable impact on the profit and loss statement². McKinsey's numbers tell a similar story from another angle: 88% of organizations report using AI in at least one business function, but only 39% can point to any measurable earnings impact from it³.

What separates companies isn't whether they adopted AI. Almost everyone has by now. What separates them is whether they can actually prove the adoption was worth it. We'll walk through both sides of that question using one running example throughout: an insurance claims workflow, the same kind of use case our energy and financial services clients bring us when they're trying to figure out where AI genuinely belongs.

Know your investment before you calculate the return

Before you can talk about ROI, you need an honest number for what you're actually investing. Most cost estimates stop at the model or API bill, which leaves out most of the real cost. A fuller picture includes:

  • Model and API costs
  • Personnel time for building and reviewing
  • Infrastructure and production deployment
  • Evaluation, testing, and ongoing monitoring
  • Human review of AI output
  • Change management and retraining
  • Ongoing maintenance and evaluation updates as new cases surface

This matters more for AI than it did for the SaaS tools that came before it. A conventional software license scales in steps. You buy another seat, another server tier, and the cost jumps up occasionally. AI processing cost scales with volume instead, climbing steadily as usage grows. Every claim your AI system triages, every document it reviews, every query it answers adds a small amount of cost. A tool that looked cheap during the pilot can look very different once it's running at full production volume, and that shift catches finance teams off guard more often than any other line item in an AI budget.

A graph illustrating the increasing costs associated with climbing a mountain over time.

Build a unit of measurement for AI work

Vehicles have a standard efficiency measure: kilometers per liter. It lets you compare a sedan and an SUV on the same terms, no matter how far either one actually drives. AI initiatives need something similar, and most companies haven't built it yet.

For AI, this unit is the cost of completing one piece of work. This could mean the cost of processing one claim, screening one resume, reviewing one document, or executing one NDA.

Back to the insurance example. Say an AI-assisted triage step costs $0.50 per claim once you include the model calls, review time, and a share of the infrastructure. At 10,000 claims a year, that's $5,000. At 100,000 claims, it's $50,000, growing in a straight line as volume grows. A COO who budgeted based on pilot-scale volume is going to be surprised by what the production invoice actually looks like.

This is especially relevant right now because manual, paper-heavy processes still govern 70% of property insurance claims⁴, so most insurers calculating AI ROI are comparing it against a real, expensive baseline rather than a hypothetical one. McKinsey estimates more than half of current claims activity could be automated by 2030⁵. Knowing your cost per claim today is what lets you track whether you're actually moving toward that number, or just adding a new expense line next to the old one.

Define what success means before you measure it

Cost per claim tells you what AI costs. It doesn't tell you whether that cost was worth paying, and answering that requires deciding what success looks like before launch, not after.

For the claims example, success could mean:

  • Reduced processing cost
  • Faster cycle time
  • Better accuracy on complex claims
  • Higher adjuster throughput without adding headcount

Pick one or two priorities that map to something the business actually cares about, rather than trying to track everything at once.

This is also where ROI and adoption split into two separate questions. An initiative can look great on paper and still fail to deliver if adjusters route around the tool, override it constantly, or only use it when it's convenient. The financial potential of AI and the behavioral change needed to capture it are two different things, and treating them as the same question is where a lot of ROI reporting falls apart.

Establish your baseline before you automate

You can't claim an improvement if you don't know what you're improving on. Before AI enters the claims workflow, it's worth documenting:

Baseline factorWhat to capture
VolumeClaims processed per month
TimeAverage cycle time from filing to decision
PeopleAdjusters and reviewers involved per claim
CostFully loaded cost per claim today
Exception rateShare of claims that don't follow the standard flow
Outcome qualityAccuracy, appeal rate, customer satisfaction

Some claims should probably stay manual. Complex or ambiguous cases often need a person's judgment more than they need speed, and pushing full automation onto that segment tends to create rework rather than savings. The baseline tells you which segment is which, and it's also what every later ROI claim gets measured against.

Usage is not adoption: measure behavior change

Usage logs tell you a system got clicked on. They don't tell you anyone's working pattern or process actually changed because of it, and that gap is where a lot of AI reporting stalls without anyone noticing.

For the insurance team:

  • Usage looks like how many adjusters log into the AI tool, and how often.
  • Adoption looks like whether they're actually resolving straightforward claims faster because of it, and whether they trust it enough to stop double checking every output.

You'll see the same pattern show up somewhere completely different. Engineering teams using AI coding assistants often watch token usage climb steadily month over month. That number alone says nothing about whether the team is shipping more features. What matters is whether AI changed how developers actually work in a way that produced more throughput, not just consuming more tokens along the way. Usage is the easy number to report. Behavior change is the one that actually predicts ROI.

Why hours saved does not equal AI ROI

Time saved is usually the first metric anyone reaches for, and it's a real number. It just isn't the same thing as value created.

Take content production. If AI cuts the time to draft an article from four hours to one, that's a genuine efficiency gain. Whether it turns into ROI depends on the resulting content performing as well as what took four hours to write. If faster output means weaker engagement, the time you saved on one side of the ledger gets spent on the other.

The claims example runs the same way. If AI triages a claim in minutes instead of hours but routes it incorrectly, the business pays for that later through an appeal, a compliance flag, or a customer who leaves after a bad experience. What actually matters is the full chain: time saved, then output, then quality, then business result. Time saved is just the first link in that chain, not the whole thing.

Measure AI ROI at three points in time, not once

A single measurement taken right after launch tells you almost nothing about whether an AI initiative earns its cost over time. It's worth checking in at three points instead.

Week one: initial usage, task completion rate, early friction points, obvious failure patterns.

Month three: whether usage is turning into changed workflows, whether early friction has actually been resolved, early signals on throughput and quality.

Month six and beyond: financial impact, operating efficiency, customer outcomes, and the real cost per claim at production scale rather than pilot scale.

Insurance claims volume is seasonal. Catastrophic weather events, open enrollment periods, and regional risk patterns all create demand spikes that a short measurement window will read as either a triumph or a failure when it's really neither. Give it enough time to see the actual pattern.

Diagram illustrating various data types for measuring a deep learning model's performance and effectiveness.

Bring investment and return together

Everything above feeds into one comparison. The investment side covers build, production deployment, ongoing operation, evaluation, and maintenance. The return side covers cost reduction, throughput, quality, productivity, and customer outcomes. Both get measured against the baseline you documented before AI entered the picture.

For the insurance example, that comparison might look like this once the initiative has run long enough to produce real numbers:

Insurers using AI-driven automation report roughly a 30% reduction in operational claims cost⁶, which is a useful benchmark for what a working initiative looks like at this stage. It's not a guarantee.

It's worth keeping this separate from AI quality evaluation. Whether an AI output is accurate enough to trust is its own production discipline, and it's a topic we've covered in our guide to identifying AI opportunities worth building. ROI measurement assumes the quality question is already handled and asks a different one: given that the output can be trusted, did the business come out ahead.

A simple web chart plotting investment against return across these dimensions, before and after AI, tends to make this comparison much easier to present to a board than a table does. Happy to put one together as a next step if that's useful.

AI ROI is a decision tool, not a report card

RAND's research on enterprise AI puts a fine point on how this usually plays out. Among AI projects that fail to deliver value, roughly a third get abandoned before reaching production, another quarter reach production but never deliver the promised outcome, and the rest run for a while without ever recovering the investment⁷. Almost none of that traces back to the model itself. It traces back to skipping one of the steps above.

The goal was never a perfect ROI number. It's knowing, with reasonable confidence, whether the value AI creates justifies what it costs to build, run, and scale. That's a lower bar than perfect measurement, and it's one any mid-market operations team can clear with the discipline this piece walks through.

Ready to know whether your AI initiative is actually paying for itself? Schedule time with our team for a structured look at your AI investment and return, before the next budget cycle asks the question for you.

FAQs

We've been running our AI tool for six months and usage is high. Isn't that a good sign?

High usage tells you people are opening the tool, not that it's changing outcomes. Check whether cycle time, cost per unit of work, or error rates have actually moved since before AI was introduced. If those numbers haven't shifted, usage alone won't show up as ROI.

How do we calculate ROI when our AI system touches multiple workflows at once?

Break it apart by workflow and calculate cost per unit of work separately for each one. A single blended ROI number across unrelated workflows tends to hide the workflow that's actually losing money.

Our AI pilot showed strong results in the first month. Can we report that as our ROI?

Treat it as an early signal, not a final number. Seasonal demand, novelty effects, and small sample sizes all inflate first-month results. Wait for the month-three and month-six checkpoints before presenting a number to leadership.

What if the ROI math doesn't work out for a workflow we've already invested in?

That's a useful outcome, not a failed one. Knowing a workflow doesn't justify further AI investment protects you from sinking more budget into it, and the baseline data you gathered is reusable for the next candidate workflow.

Should every workflow be measured with the same ROI framework?

No. High-volume, standardized workflows like claims triage suit the cost-per-unit approach well. Judgment-heavy, low-volume workflows are often better measured by quality and risk reduction than by cost per unit, since the volume needed for a clean unit cost isn't there.

Authors

Sarthak Dudhara

CEO & Co Founder
Co-Founder and Chief Technology Officer at Aubergine. Firm believer in "actions speak louder than words". There is nothing that gets me as excited as building new and exciting things that disrupt the status quo.

Podcast Transcript

Episode
 - 
8
minutes

Host

No items found.

Guests

No items found.

Have a project in mind?

Read