
Start here
A consulting team hands a client memo to an assistant and expects a uniform boost. The best evidence says the boost depends on which side of an invisible line the task sits on. A pre-registered field experiment with Boston Consulting Group, described in the working paper, assigned 758 BCG consultants, about 7 percent of the firm's individual-contributor consultants, to one of three conditions: no AI, GPT-4 access, or GPT-4 access plus a prompt-engineering overview. Deciding whether to lean on an assistant for a deliverable is really asking whether it sits inside or outside what the authors call the 'jagged frontier.'
What the documents say
The design was a randomized field experiment, not a survey: consultants were randomly split after a baseline task established starting skill. Across 18 tasks chosen to be within GPT-4's reach, drafting a press release, segmenting a market, generating product ideas, AI users completed 12.2 percent more tasks, finished 25.1 percent faster, and produced work rated over 40 percent higher in quality than the control group, per the paper's numbers. Gains were uneven: below-average consultants improved 43 percent, above-average consultants 17 percent. On a task chosen to sit outside the frontier, the pattern reversed: AI users were 19 percentage points less likely to reach a correct solution. The Harvard institute's summary repeats the inside-frontier figures but omits the decline, a gap worth noticing when a summary is all that gets read. BCG's own publication frames the work as academic partners analyzing BCG's data, not a BCG-authored study; funding came in part from Harvard Business School.
Check this
A reader can check which side of the frontier a task sits on by asking whether it resembles the 18 bounded, rubric-judged tasks used, or a genuinely novel judgment call picked to be hard. The frontier is not about AI being smarter some days; it reflects which problems appear often enough in training to be done reliably, versus ones where a confident answer can be wrong. Testing a small piece against a known answer before trusting a large piece is what the design effectively required.
What holds and what fails
The gains hold most convincingly for well-defined, judged tasks like the ones tested, on one firm's consultants using one 2023 model, a study that itself calls the frontier 'jagged' rather than fixed. It fails, by the paper's own result, when a task looks similarly difficult but sits past the model's reliable range, and the paper found no way for consultants to sense that boundary from inside a task. Treating one study as proof for a different company or year is editorial overreach the paper does not make.
- Before delegating a task, ask whether it resembles a graded task or a genuinely novel judgment call.
- Read a working paper's abstract directly rather than a summary that may drop an inconvenient finding.
- Treat a confident output on an unfamiliar task as unverified until checked against a known answer.
The study's real contribution is not one productivity number; it is evidence the same tool can help and hurt within one afternoon of work, depending on the task handed to it.
Sources & reading trail
States the pre-registered design, 758-consultant sample, in-frontier productivity and quality gains, the outside-frontier accuracy decline, and Harvard Business School funding.
Source published: 22 September 2023 · Retrieved: 16 September 2026
Institute summary confirms sample size and in-frontier speed, quality and completion figures, and illustrates how a secondary summary can omit the outside-frontier finding.
Source published: Not established · Retrieved: 16 September 2026
BCG's own framing of the collaboration as academic partners analyzing BCG's data, confirming the 758-consultant sample and BCG's role as research partner rather than study author.
Source published: 21 September 2023 · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.