# How to measure the ROI of AI in operations

> AI operations ROI measured as cost per operated process: where agents pay off in high-volume back-office work and why pilot numbers do not transfer.

- Canonical: https://arkatai.com/en/ai-operations-roi/
- Site: Arkatai (https://arkatai.com) — agentic operations as a service
- Language: en
- Published: 2026-07-18

---
When a slide says "40% productivity gain," ask what was divided by what. Without a denominator and baseline, the percentage is not an ROI measurement. For agents in operations I use one unit: cost per process operated. That unit lets a company compare a tool, an internal build and a managed operation on the same basis.

## The unit that matters: cost per operated process

"Productivity" is insufficient when nobody in the room can reproduce the calculation. A unit tied to a process can be checked against company records.

The unit I use is the cost of operating one process, end to end, before and after. Take a single process, say supplier invoice handling, and price what it costs you today: the hours of human work per invoice, the cost of the errors that slip through, the cost of the delay when an invoice sits in someone's inbox for six days, the cost of the person who does nothing but chase the other people. That number exists in your company right now. Almost nobody has computed it, which is itself informative.

Then price the operated version: the service fee or amortized internal build, model spend per case, human minutes still required for review and exceptions, and ongoing maintenance. Divide by volume in both worlds and you have two comparable numbers, each one auditable, each one arguable in a way a productivity percentage never is.

That comparison exposes the volume, labor, error and maintenance assumptions behind the result.

## The ROI appears in high-volume processes

The processes that often produce a measurable result are back-office operations with repeated cases.

Invoice matching. Order intake that arrives as free-text emails and PDFs. Keeping supplier and product master data coherent across systems. First-pass triage of claims and returns. Reconciliations at month-end. These processes share a profile: high volume, mostly repetitive, rules that are largely knowable, and errors that are visible and priceable. That profile is precisely what makes the denominator large and the measurement clean. An agent that removes twenty human minutes from a case you handle nine thousand times a year is a number your CFO can check.

A customer-facing assistant may touch revenue, but attribution is harder. Did the assistant close the sale or did the discount? If the effect cannot be separated, the company cannot assign a return to the assistant with confidence.

Back-office processes also contain many company-specific exceptions and unwritten rules. Encoding them is part of putting [agents into operations](/en/ai-agents-in-operations/) and is [custom-phase work](/en/the-custom-phase/) because a product does not arrive with that knowledge.

## The pilot is not the operation

A common measurement error is applying a number from the pilot to the production operation.

A pilot often uses curated cases, close supervision and immediate vendor support. Production includes malformed invoices, month-end peaks, holiday handovers and inputs absent from the pilot. Annualizing pilot numbers assumes those additional costs and exceptions do not exist.

I discuss this gap in [why enterprise AI pilots fail](/en/why-enterprise-ai-pilots-fail/). For ROI, measure after the agent has covered a complete business cycle, including peaks, absences and exception-heavy customers. Earlier numbers are useful for engineering decisions but do not represent steady operating cost.

## What goes into the spreadsheet

The spreadsheet should contain the following inputs.

Before deployment, count cases per month, human minutes per case, observable error rate and cycle time from arrival to completion. Record the confidence and gaps in the baseline rather than omitting it. If it cannot be established, the company has a readiness problem, covered in [preparing your company for agents](/en/prepare-your-company-for-agents/).

After the agent operates the process, count the same things, plus three that are new: minutes of human review per case, escalation rate to humans, and the full recurring cost of the service or internal system. Review time and escalations belong in the cost, always. An agent that resolves seven out of ten cases and hands three to a person is a fine economic object, but only if you price the three.

Measure for a couple of quarters before scaling the conclusion. During early operation, exceptions surface and the escalation rate may fall as rules are added. Use a defined review period and report the trend instead of treating the first month as steady state.

## Who owns the measurement owns the number

The sourcing decision also determines which instrumentation the company controls.

When a process runs inside a product, its dashboard defines what is counted. Escalation time, workarounds and corrections outside the product may be missing because the system cannot observe them. A managed operation should instead log each action against the client's metric and expose the inputs behind the result. The sourcing choices are compared in [buy, build or contract the outcome](/en/buy-vs-build-enterprise-ai/).

## Questions boards ask me

### What payback period should we demand?

Require a measured baseline and dated review before setting a payback period. A period quoted without volumes and exception rates is an estimate based on assumptions. The sequence is baseline, steady-state measurement and then payback arithmetic.

### Should we start with the process where the ROI is biggest?

Start where the ROI is measurable and individual errors are tolerable. The first process teaches the organization how to review, escalate and govern the service. Those controls can then be reused in processes with greater financial impact.

### What about revenue-side AI? The upside seems larger.

Revenue-side upside may be larger, but attribution is often weaker. Back-office returns can be smaller and still verifiable per case. The two categories can proceed with different confidence levels and measurement methods.

### Our vendor showed us a customer with strong results. Why shouldn't we expect the same?

That customer's number depends on its volumes, exception profile and process discipline. A reference shows that the mechanism has worked under those conditions. It does not price the same mechanism in your operation. Your baseline does.