01

What changed

AI-at-work studies often collapse unlike tasks into one average. This experiment was designed to expose the boundary. Researchers recruited 758 Boston Consulting Group consultants and first gave everyone a baseline task without AI. Participants were then assigned to one of two separate task arms and randomly placed in one of three conditions: no AI, access to a custom GPT-4 interface, or the same GPT-4 interface plus a prompt-engineering overview.

In the 385-person arm designed around work GPT-4 could handle, consultants developed a new footwear product through 18 subtasks involving creativity, analysis, persuasion, and writing. The no-AI group completed 82% of the subtasks. The GPT-only and GPT-plus-overview groups completed 91% and 93%. AI users reached the final question more than 25% faster on average, and their human-graded response quality was 32% higher. Separate footwear-design graders reproduced the creative-idea advantage.

The other 373 consultants faced a business case that required combining a spreadsheet with interview notes containing a crucial contradiction. The spreadsheet alone pointed toward the wrong brand, and GPT-4 was prone to follow it. The control group reached the correct recommendation 84.5% of the time, versus 70.6% with GPT-4 alone and 60% with GPT-4 plus the overview—an average 19-percentage-point drop across the AI groups. Those answers arrived faster and sounded more coherent, even when they were wrong.

02

What this could change for you

The practical lesson is to classify the task before reaching for the tool. Drafting options, reorganizing material, generating variations, and producing a first-pass synthesis can be good candidates when you can inspect the output. A recommendation that depends on finding one disconfirming fact across several sources needs a different workflow: read the decisive evidence yourself, state the failure condition in advance, and verify the conclusion independently.

For teams, an AI policy should be organized around tasks rather than job titles. A consultant, marketer, analyst, or manager may cross the helpful and harmful sides of the boundary in one afternoon. Keep an explicit list of uses that have been tested, the source material required, what a reviewer must check, and what outcome would reveal that the model failed. Update the list when the model or workflow changes.

Prompt training alone is not a safety system. In this experiment, the overview improved quality on the suitable tasks, but the overview group performed worst on the deceptive case. Knowing how to get a polished answer is different from knowing whether the model should answer at all. Verification matters most when the output is persuasive, the evidence is scattered, and a wrong answer can still look complete.

03

What it does not prove

The participants were early-career consultants from one elite global firm, and the work was a time-boxed experiment rather than a live client engagement. The tasks were realistic and designed with firm leaders, but they did not measure client outcomes, long-term learning, worker well-being, pay, employment, confidentiality risk, or whether managers could build a durable review process.

The system was GPT-4 as available in April 2023. Current models, tools with retrieval, and longer-context workflows may move the boundary, but the experiment did not test them. It also tested one broad creative-product exercise and one deliberately difficult business case; it cannot estimate what share of ordinary knowledge work sits on either side of the frontier.

Finally, the paper's 'jagged frontier' framing was not itself preregistered. Researchers selected tasks intended to reveal strength and weakness, and quality depended on expert rubrics rather than market performance. The study strongly demonstrates that the same professionals can improve on one task and deteriorate on another. It does not supply a permanent map of which tasks are safe to delegate.

The bottom line

AI made consultants faster and better when the task matched the model's capabilities, then made them faster, more persuasive, and substantially less accurate when a crucial fact contradicted the obvious pattern. The useful rule is not simply to use AI or avoid it. Use it where the output can be checked, test each workflow against known answers, and reserve independent human judgment for conclusions that depend on reconciling conflicting evidence.

Primary research

Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

Organization Science · 2026 · DOI 10.1287/orsc.2025.21838

View the research ↗