Essay · Implementation

Why do enterprise AI pilots fail?

Arkatai 4 min

A pilot can meet its demonstration criteria while no company process changes eighteen months later. I have seen that outcome with different vendors, use cases and budgets. Four design errors recur, often in combination. Model quality alone rarely explains the result.

The category error: buying product-phase for a custom-phase problem

The first error occurs when the company classifies what it is buying.

The stack contains components in different phases. Models are utility and orchestration tooling is consolidating into product. Encoding a company’s processes, exceptions and unwritten rules remains custom because the content differs by company. The custom phase explains why the market may price the whole stack as product even when implementation still includes that custom work.

The platform may deliver the utility and product layers it promised. The pilot then reaches the custom layer, where someone must work with operations to encode accumulated exceptions. If nobody budgeted or owns that work, the project stops between demonstration and production. From the steering committee, this can look like a technology failure even though the missing scope was organizational and implementation work.

No process owner, no process

The second error is assigning a project manager but no process owner.

A project manager delivers scope and schedule. A process owner can change the process, accept the risk of agent actions and resolve differences between the manual and current practice. A pilot in an innovation team or IT sandbox may meet technical criteria while the operating team has no authority, incentive or obligation to adopt it.

Ask who is accountable if process performance declines next quarter because of the agent. A committee may set policy, but the answer also needs a named process owner who accepted the pilot. AI agents in operations describes the same requirements—ownership, permissions and limits—at production scale.

Prepared data in the demo, operational data in production

The third failure is the gap between the data the pilot was fed and the data the operation produces.

Demos use curated inputs: well-formed invoices, expected ticket formats and matching master data. Production includes photographed spreadsheets, duplicate customer records and fields filled with “see notes.” These are common consequences of years of system and process changes.

Handling company-specific data requires identifying each recurring case, deciding the expected behavior and encoding it. A pilot that excludes those cases tests only a selected subset of the process. Writing down the process with its exceptions produces the operating model as code and creates an asset that later systems can reuse.

Success defined as “it works” instead of “it operates”

The fourth error is in the success criteria.

“It works” can mean that the agent produced correct outputs on test cases. “It operates” requires representative data and volume, defined permissions and limits, a working escalation path, review of exceptions and a measurement period long enough to compare costs. Each condition must have an owner.

A pilot defined only by test output leaves production work for a second project: integration, exceptions, ownership, controls and maintenance. If that work was absent from the original scope and budget, scaling is not an extension. The sourcing decision must include a third option beyond buy or build: contract an operated service whose provider implements and maintains the company-specific layer.

At kickoff, define success as a named process operating under its owner for an agreed period and meeting specified metrics. That definition also identifies pilots that should not start.

The gate from pilot to operation

Before approving production, the committee should receive evidence for four decisions. The process owner confirms which steps the agent replaces and which remain with people. Technology confirms the identities, permissions, limits and recovery procedure. Operations reports volume, escalation rate, error classes and the time spent reviewing cases over a representative cycle. Finance compares the cost per completed case with the baseline, including maintenance and human review.

Each item needs a threshold agreed before the measurement period. A statement such as “accuracy was acceptable” is not a threshold. “Fewer than two incorrect classifications per thousand cases, with no unauthorized write action” is. The decision can then be expand, continue at the same scope, revise or stop. All four outcomes are valid if they follow the evidence. What should not happen is a fifth outcome in which the pilot remains nominally active while no team owns its next decision.

Questions boards ask me

Our pilot succeeded technically but went nowhere. Restart or salvage?

Often it can be salvaged. Take the use case to the executive who owns the process. If that person accepts it, re-scope around representative volumes, exceptions and operating criteria. If no process owner wants it, close the project and record that decision.

How long should a pilot take?

Long enough to cover a full cycle, including month-end, campaign peaks or seasonal load where relevant. A shorter test may omit the cases that determine support and exception costs. Model integration can be quick. Documenting and encoding the process usually takes more calendar time.

Should we run several pilots in parallel to raise the odds?

Parallel pilots compete for the attention of process owners and people who hold undocumented knowledge. If that capacity is limited, take one pilot through production before starting several. It produces the first codified process and a tested set of governance controls.

How do we tell before signing whether a vendor pilot will hit the custom layer?

Ask what they need from the client. If the answer is mainly data and a sandbox, ask who will document exceptions and assign the process owner. A proposal should include time with operators, exception discovery and named ownership before implementation starts.