01

What changed

Most AI coding stories still rely on demos, benchmark puzzles, or developer self-reports. This study asked a harder question: what happens when large companies randomly give some developers access to an AI coding assistant and compare real work output against colleagues who did not get the tool yet? The paper pools three field experiments run in the ordinary course of business at Microsoft, Accenture, and an anonymous Fortune 100 company.

The intervention was narrow. Developers in the treatment groups got access to GitHub Copilot, an AI coding assistant that suggests code completions inside normal development tools. The study did not force people to use it, and the authors stress that the treatment was access, not guaranteed adoption. That makes the claim about workplace rollout more believable and more limited at the same time.

Across the pooled experiments, the authors estimate that access to the tool increased completed tasks by 26.08%, with a standard error of 10.3%. Secondary workflow measures moved in the same direction: weekly commits rose 13.55% and weekly builds rose 38.38% in the pooled estimates. The paper also reports that less-experienced developers adopted the tool more and saw larger productivity gains.

02

What this could change for you

For ordinary users of software, the believable upside is not that one tool solved programming. It is that some product teams may ship routine fixes, integrations, and internal development work faster when the assistant fits the job. That could show up as shorter waits for minor product improvements, bug fixes, or internal tools employees rely on every day.

For managers and teams, the study offers a grounded way to think about workplace AI. The strongest measured benefit was not universal genius or autonomous software creation. It was more completed development tasks in real organizations after a randomized rollout, with bigger help for developers who had less experience and higher willingness to adopt the tool.

The broader lesson is that low-risk AI gains often come from narrow workflow assistance rather than full automation. The assistant did not replace product decisions, code review, testing, or team accountability. It sat inside an existing software workflow and helped some developers get more units of work across the line.

03

What it does not prove

The paper does not prove that AI-generated code was better code. Completed tasks, commits, and builds are workflow outputs, not a direct measure of software quality, maintainability, security, or downstream customer satisfaction. The build-success measure did not show a clear pooled improvement, so the honest claim is about task completion rather than code quality.

The pooled estimate hides noisy and uneven company-level results. The three experiments had different designs, durations, and adoption patterns, and the authors say statistical power remained a challenge because developer output is highly variable and control groups eventually got access. This is stronger than a lab study, but it is not a claim that every software team should expect the same 26.08% gain.

The study also cannot settle the bigger labor questions people actually care about. It does not prove higher wages, better jobs, fewer jobs, faster promotions, or long-term productivity across all knowledge work. It measures a bounded workplace effect in software development after randomized access to one coding assistant.

The bottom line

This is one of the clearest routine AI workplace results available because it tests real developers in real companies with randomized access rather than a benchmark or a survey. The measured gain is not magical and it is not universal. But a three-company field study finding more completed developer tasks after AI-assistant access is a strong, low-risk everyday-work result, especially because the paper stays honest about uneven adoption, noisy company-level estimates, and unresolved code-quality questions.

Primary research

The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers

Management Science · 2026 · DOI 10.1287/mnsc.2025.00535

View the research ↗