What changed
Most AI-at-work research tests a tool against no tool. This field experiment tested the surrounding job design. A multinational pharmaceutical company introduced the same machine-learning sales system across 12 business units in Denmark, Finland, Norway, and Sweden. Researchers randomly assigned the units to three conditions: a legacy information system, the new AI system with work settings that did not match employees' cognitive styles, or the same AI system with procedures, decision authority, training, and incentives tailored to those styles.
The company tracked 72 sales experts for 22 quarters: nine before the intervention and 13 after it, producing 1,584 person-quarter observations. Compared with the legacy-system group, the untailored AI group conducted 0.46 fewer daily sales meetings after rollout, about a 20% decline. The tailored group conducted 0.84 more daily meetings, an increase of more than 40%. These meeting effects held across the paper's clustered and randomization-based checks.
Units sold moved in the same direction. The untailored group sold roughly 547 fewer units per person per day than control, while the tailored group's estimate was about 450 more. But that positive tailored estimate was not statistically significant in the main individual-level model. The strongest defensible finding is therefore about meetings—and about how implementation shaped whether the same AI system helped or hurt.
What this could change for you
If AI is arriving in your job, the practical question is not only whether the model is accurate. Ask what else is changing around it. Does the tool force a rigid process on work that depends on judgment? Can you override its recommendation? Does training reflect how you solve problems? Do incentives reward useful decisions or merely system usage? In this experiment, those surrounding choices separated a performance gain from a performance loss.
For managers, the study argues for treating AI rollout as job design rather than software installation. The researchers classified employees as more adaptive or more innovative thinkers, then aligned four organizational settings to those preferences. The exact questionnaire and split are not a plug-and-play management recipe. The transferable lesson is to test workflows with the people doing the work, preserve appropriate authority, and measure real output—not logins alone.
The system itself was identical in the two AI groups. Usage later diverged: daily logins declined in the untailored condition and rose in the tailored condition. A mediation analysis linked some of the meeting difference to usage, but it could not rule out reverse causality. Workers may have used the system less because the surrounding design was already hurting their performance.
What it does not prove
This was a small cluster experiment: 72 people, 12 randomized business units, and only four units per condition. The paper reports low statistical power and explicitly cautions against strong causal interpretations despite consistent sensitivity analyses. Baseline performance also differed across groups, although the authors used difference-in-differences models, individual fixed effects, time controls, and several robustness checks.
The setting was one pharmaceutical company's Nordic sales operation, and the system ran from 2012 to 2017. It was a machine-learning sales tool, not a modern generative-AI assistant. Meetings and units sold are commercial performance measures; the study did not measure worker well-being, job satisfaction, product quality, patient outcomes, layoffs, or the cost of tailoring.
Finally, the intervention changed four things together: procedures, authority, training, and incentives. The experiment cannot isolate which one mattered most, nor cleanly separate the effect of those work changes from the AI interaction itself. It supports designing the human system around AI carefully; it does not validate personality testing as a universal deployment method.
The bottom line
The same AI tool helped one group and hurt another depending on how the work around it was designed. In this long-running field experiment, a tailored rollout increased daily sales meetings by more than 40% versus control, while an untailored rollout reduced them by about 20%. The result is unusually concrete but power-constrained and narrow. Before judging an AI system by its demo, test whether the procedures, authority, training, and incentives around it fit the people expected to use it.
Primary research
Human-Centered Artificial Intelligence: A Field Experiment
Management Science · 2026 · DOI 10.1287/mnsc.2022.03849


