Sofpact field note

AI in Service Operations: Where the Measured Gains Land

Editorial bar chart showing the AI productivity gain concentrated among the newest staff

Most service operations leaders have now sat through the same demonstration. An assistant drafts a reply, summarises a case history, proposes the next step. It looks convincing. The harder question arrives straight afterwards: if this goes in front of two hundred people on Monday, what actually changes, and for whom?

There is a published answer to that, and it is unusually good evidence. It is worth knowing what it says before committing an operation to a rollout, because the shape of the result is not the one most business cases assume.

What the largest published study found

Brynjolfsson, Li and Raymond: Generative AI at Work, published in The Quarterly Journal of Economics in 2025, followed 5,179 customer support agents at a software firm as access to a generative AI assistant was released in stages. The staggered release is what makes it worth reading: agents with the tool and agents without it were doing the same work at the same time, so the comparison is a measurement rather than an impression.

Three findings stand out.

  • Productivity rose by around 14%, measured as issues resolved per hour.
  • The gain concentrated among newer and lower-performing agents, at roughly 34%. For the most experienced agents it was close to nothing.
  • Customer sentiment improved and staff retention rose — two things that seldom move together in a service operation, and neither of which was the tool’s stated purpose.

The authors’ explanation is the useful part. The assistant appeared to spread the working practice of the strongest agents to everyone else. It was not making people quicker at what they already knew. It was closing the distance between a new starter and an experienced one.

The shape of the gain matters more than its size

An average of 14% is the figure that reaches the business case. It is also the least useful figure for planning, because almost nobody in the operation experiences it. Some people gain a third. Some gain nothing.

Where the value actually sits

If the benefit tracks inexperience, its size in a given team is a function of how much experience variation that team contains. A settled team of long-tenured specialists handling familiar work has little headroom. A team with seasonal intake, recent growth, high turnover or a broad and changing product surface has a great deal. The same tool is worth materially different amounts in two teams of identical size.

That is a better selection rule than the usual one, which is to start wherever someone volunteered.

What to measure

Issues resolved per hour is a legitimate measure inside a controlled study. Inside a live operation it is weak on its own, because the quickest way to resolve more issues per hour is to resolve them worse. Pair any throughput measure with at least one quality measure produced independently of the people being measured: reopened cases, escalation rate, complaint rate, quality sampling, or first-contact resolution confirmed after the fact.

The NIST AI Risk Management Framework treats measurement as an activity running through design, deployment and use rather than a gate passed once. In a service operation that translates plainly: the evaluation that justified the rollout should still be running six months later, on live work.

What happens to your experienced people

If the tool works by distributing what good agents already do, then good agents are the source of the standard rather than the beneficiaries of it. That is a role, and it deserves to be named and resourced — reviewing suggested responses, correcting the knowledge the assistant draws on, and deciding which cases should never be handled with assistance at all. An operation that gives its best people no part in the system will watch the quality of suggestions drift towards the average of whatever the model was given.

What happens to training

A tool that compresses the experience curve does not remove the need to build experience; it changes when that need is felt. New agents reach acceptable output sooner and deep understanding later, because the intermediate struggle that used to produce it has been shortened. Operations that care about their five-year bench, and not only their five-week ramp, should plan for that rather than discover it.

The controls that make it a deployment rather than an experiment

Service work touches customers, and in regulated sectors it touches obligations. The Bank of England and FCA survey of AI in UK financial services found 75% of firms already using AI with a further 10% planning to within three years, and 84% with a named person accountable for their AI framework. The same survey found only 34% of firms reporting a complete understanding of the AI they use, and 46% reporting partial understanding.

That gap — near-universal adoption, majority-partial understanding — is the practical risk in a service rollout. Four controls close most of it.

  • A preserved escalation path. A customer must be able to reach a person, and an agent must be able to reject a suggestion without justifying the refusal. A suggestion that cannot be refused is a decision.
  • Grounding in approved material. The assistant should draw on current, owned, versioned content. A confident answer citing a withdrawn policy is worse than no answer.
  • A record of what was suggested. When a complaint arrives eight months later, the question will be what the agent was shown at the time. That is answerable only if it was logged then.
  • A stop condition agreed in advance. Name the measure and the threshold that would pause the rollout before the rollout starts, while the decision is still cheap to make.

The Artificial Intelligence Playbook for the UK Government makes a related point about evidence: AI systems generally warrant more substantial evaluation than other interventions, and a staged or controlled comparison is how that evaluation is obtained. That is exactly the design that made the study above worth reading, and it is available to any operation willing to release in waves instead of all at once.

A first deployment worth running

Choose one queue rather than one department. Pick work with a clear definition of done and an observable quality signal. Release to half the team first and hold the other half as the comparison for a defined period; staged release costs nothing extra and is the difference between knowing and believing.

Agree in advance what would count as success, what would count as harm, and who decides. Record the baseline before the tool arrives, because it cannot be reconstructed afterwards. Then run it long enough for the novelty to wear off — the first fortnight measures enthusiasm.

Adoption on its own is no longer a differentiator. The Stanford HAI 2026 AI Index Report puts organisational adoption at 88%, while noting that results remain uneven and that documented AI incidents rose to 362 in 2025 from 233 the year before. Nearly everyone has deployed something. The operations getting a return are the ones that can say what changed, for whom, and how they know.

If you want an independent read on a service workflow before committing to a rollout, customer operations work usually starts with a scoped assessment.