01

What changed

The usual AI-at-work story treats the model as a solo assistant: a faster search box, a better autocomplete, or a chatbot that helps one person finish a task sooner. This paper asked a more realistic question: what happens when AI is not just a tool on the side, but a teammate inside a shared workspace where both sides can write, edit, choose images, and shape a finished output together?

To test that, the researchers built Pairit, a collaboration platform designed to log every message, edit, and delegated action. They recruited 2,234 participants through Prolific, matched them into either human-human or human-AI teams, and assigned them the same task: create ad campaigns for a think tank's annual report. The teams produced 11,024 ads in total, along with 182,607 messages, 1,889,559 text edits, 62,119 image edits, and 10,074 AI-generated images.

The productivity gain was real. Human-AI teams produced 50% more ads per worker than human-human teams. But the quality result was split rather than uniformly better. Human-AI teams produced stronger ad text, while human-human teams produced stronger image choices. That is the paper's practical claim: AI helped where drafting, rewriting, and following instructions mattered most, but it lagged where aesthetic judgment and visual selection mattered more.

The authors then ran the ads in a real X campaign that generated more than 4.9 million impressions. The field test matched the quality split. Better text was associated with higher click-through rates and longer viewing duration after the click, while better images were associated with lower cost per click. AI changed output quality in ways that mattered downstream, but the gains and losses landed in different parts of the funnel.

02

What this could change for you

For people already using AI at work, the useful takeaway is narrower than "let the model do the project." AI looked strongest here as a draft-heavy partner: generate options, rewrite copy, suggest structure, and handle repetitive production work. Humans still looked better at choosing the most compelling image and judging the final mix.

The teamwork pattern matters too. Human-AI teams sent 25% more task-oriented messages, delegated 17% more work, and made 62% fewer direct text edits. That suggests the skill that matters is not only prompting. It is learning how to run a clearer, more directive workflow: give the system specific tasks, review the output fast, and keep moving instead of polishing every line manually.

There is also a cost. The more work teams delegated to AI, the more homogeneous their outputs became. If your job rewards one solid first draft, that may be a worthwhile trade. If it rewards originality or a wide range of creative options, heavy AI delegation can quietly narrow the result even while average output improves.

03

What it does not prove

This was not a study of long-running office teams inside a company. The participants were recruited through Prolific, the collaboration happened in a newly built experimental platform, and the task was ad creation for a think tank. That is more realistic than a benchmark, but it is not the same as following marketers, product teams, or agencies through months of real work.

The paper is a preprint. The random assignment is strong, and the field test adds real-world validation, but the mechanism claims around recognition, communication style, and delegation are partly post hoc and correlational. The authors explicitly say they cannot fully separate whether people behaved differently because the partner was AI or because they recognized it as AI and adjusted.

The result is task-specific. Human-AI teams underperformed on image quality, and the authors warn that their findings may generalize best to multimodal creative tasks with similar text-versus-visual tradeoffs. The paper does not show that AI teammates are ready for high-stakes decisions, deep expert domains, or situations where a wrong output is costly and hard to reverse.

The bottom line

AI did not solve teamwork in general. It solved a narrower and more practical problem: producing more draft work, faster, in a shared creative workflow. In this randomized experiment, the best case for an AI teammate was not better judgment across the board. It was better throughput on text-heavy work, with humans still needed for visual judgment, quality control, and protecting against sameness.

Primary research

Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance

arXiv preprint · 2026 · DOI 10.48550/arXiv.2503.18238

View the research ↗