
Imagine finishing a proposal in 20 minutes instead of an hour. You send it for approval. It sits in the same queue until Thursday.
You have saved 40 minutes. The client still gets an answer on Thursday. Both things can be true, and the difference matters when a company starts counting the returns on AI.
Research on AI at work is making that distinction harder to ignore. Time spent on a task, completed work and the quality of the working day are separate outcomes. A useful AI strategy needs to say which one it is trying to improve.
A six-month randomized field experiment across 66 firms gave selected employees access to AI inside their existing email, meeting and writing applications. In the experiment’s second half, the researchers estimate that users of the tool, who represented 80% of the group offered access, spent about two fewer hours a week on email and worked less outside normal hours. The study is forthcoming in American Economic Review: Insights. Read the journal abstract.
The scope matters: that is an estimate for users, not a two-hour saving established for every employee offered a license. The researchers also did not detect changes in the quantity or composition of tasks following individual access to AI. That limits what the experiment can tell a business about additional output.
Microsoft conducted the trial using Microsoft 365 Copilot, with data collected between September 2023 and October 2024. Three of the four authors work at Microsoft Research. Those details appear in Harvard Business School’s account of the research.
Other evidence measures a more direct increase in completed work. A paper published online in Management Science on February 27, 2026, combines randomized experiments involving 4,867 developers at Microsoft, Accenture and an unnamed Fortune 100 company.
The pooled estimate was roughly 26% more completed tasks among developers using an AI coding assistant. Results differed across the experiments and were noisy; the reported standard error was 10.3 percentage points. Less experienced developers adopted the tool more and saw larger gains. The author team includes Microsoft researchers. Read the published study.
This is evidence about particular coding assistants, workers and tasks. The percentage does not establish an equivalent increase in revenue, software reliability or the output of every engineering team using today’s agents. It measures something narrower, which is precisely why it is useful.
A separate study, published in the Quarterly Journal of Economics in 2025, followed the staggered introduction of an AI assistant among 5,172 customer support agents. It found a 15% average increase in issues resolved per hour. Less experienced workers improved speed and quality; the most experienced workers saw small speed gains and small quality declines. This was a study of a rollout, rather than the same randomized design as the developer trials. Read the authors’ paper.
The practical implication is to examine gains by role and experience. A department-wide average can conceal a tool that helps a new colleague considerably and adds little for the person training them.
On February 24, 2026, research organization METR reported that its follow-up experiment on developer productivity could not provide a reliable estimate of AI’s current effect.
Some developers were unwilling to participate because tasks could be assigned to a condition where AI was prohibited. Others withheld tasks they expected AI to help with most. The researchers also cited lower participant pay and difficulty tracking time when people used several agents concurrently. These problems affected who and what the experiment could measure.
METR thought developers were probably getting more benefit than in its early-2025 study, but said the newer data offered only weak evidence about the size of that change. It began revising its approach. Read METR’s methods update.
That leaves an important limit on headline comparisons: a result belongs to a particular period and experimental setup. A number measured with earlier tools cannot settle what a different team will gain from a newer system.
Here is a practical way to apply these findings. Consider a hypothetical team producing client proposals. Before AI, drafting takes 60 minutes and review takes 20. With AI, drafting takes 20 minutes and review takes 35 because the reviewer must check additional claims and remove unnecessary material.
The team has saved 25 minutes of labor per proposal. Counting only the drafting stage would report a 40-minute saving and miss the extra work handed to the reviewer. Neither figure tells you whether the proposal reaches the client sooner.
To find that out, the team needs to follow the whole process: preparation, generation, checking, revisions and approval. A faster draft only shortens the delivery schedule if the work that follows can move sooner too.
This example also suggests a straightforward intervention. If approval is the delay, try a daily review window with a clear owner. Then check whether client responses arrive sooner while meeting the same quality standard. The useful experiment is specific enough to reveal where the benefit disappears.
For a workplace trial, choose one recurring type of work and agree on the intended outcome before introducing the tool. Four measures are a reasonable starting point:
Compare similar tasks over a defined period. Where feasible, stagger access randomly so an unusually easy month or a team of enthusiastic early adopters does not become the entire explanation for an apparent gain.
Keep the measures proportionate. A small team does not need invasive monitoring to learn whether its proposals require fewer revisions. A shared record of task times and outcomes can be enough to identify a bottleneck worth testing.
Most of all, decide what success means for the people doing the work. Finishing the same amount to the same standard with fewer evening hours is a valuable outcome. Producing more useful work can be another. Treating every saved minute as an automatic commitment to higher volume quietly makes that choice for everyone.
The next step after a successful AI trial should be a concrete decision about the time it frees up: which deadline moves forward, which backlog gets attention, or which evening becomes someone’s own again.