Two Studies, Two Opposite Answers


I noticed something in my own writing a while back. Before AI, the gap between strong and weak writers was large. Once AI became common, that gap seemed to shrink. People with strong judgment could get even better output out of AI, but their own core judgment probably hadn’t changed much. People who started weaker saw a much bigger jump — AI-assisted drafts that were clean and logically coherent in a way their own writing rarely was. I went looking for research that might confirm this and found two studies. They land on opposite answers.

An NBER study on customer service agents, by Brynjolfsson, Li, and Raymond, found that AI assistance raised agent productivity by about 15% on average, and that gain went disproportionately to newer, less-skilled workers. Their read is that generative AI may be encoding top performers’ best practices and handing them directly to less experienced staff. A separate study on writing tasks found something similar: professionals who scored lowest before AI improved far more with assistance than those who already scored highly. Both studies land on the same side — the gap narrows.

A Columbia Business School study led by David Holtz and colleagues, on Kenyan entrepreneurs using AI for business advice, found the opposite. AI didn’t narrow the gap. It widened the inequality between the best and worst performers. The reason comes down to this: in open-ended, ambiguous situations, AI doesn’t replace human judgment, it amplifies it. The value of what AI produces depends heavily on the judgment of the person steering it. Weaker performers tended to take AI’s generic advice at face value, letting it substitute for thinking they should have been doing themselves. Stronger performers stayed skeptical and used AI’s suggestions only to make changes that actually fit their situation.

These findings look contradictory on the surface, but they reconcile once you separate the tasks by how well-defined they are. Customer service scripts and first drafts have clearly bounded problems with a limited answer space — AI can copy over best practices directly, and the gap narrows. Call this the execution layer. Business strategy is different: open-ended, no fixed answer. AI hands you generic advice, and the real question is whether you can judge if that advice fits your situation. Call this the definition layer. What weaker performers lose there isn’t execution ability — it’s not knowing whether to trust what AI just told them. Compression happens at the execution layer. Divergence happens at the definition layer.

Both studies have two problems buried in them, though.

The first is the gap between a snapshot and a trend. These studies can only measure AI as it exists right now. They can’t tell you what happens once AI’s judgment keeps improving beyond today’s baseline. Using them to argue that compression is happening right now is fine. Using them to argue that compression will keep going until judgment itself gets flattened is a much bigger claim than the data can carry — that kind of long-run claim needs a precedent that’s already run its full course, like chess, rather than a cross-section still mid-arc.

The second is that the measurement scale itself might be manufacturing the result. The customer service study used resolution rates and handling time, close to a ratio scale, but there’s a ceiling-and-floor effect hiding in there: low performers start far from the ceiling, so they mathematically have more room to improve, while high performers are already near the top, so even an equal underlying gain in judgment shows up as a much smaller measured improvement. Some of “low performers improved more” might just be an artifact of the floor effect, not evidence that judgment actually flattened. Writing scored on a rubric is even messier: the psychological distance between going from 3 to 5 and going from 7 to 9 isn’t necessarily equal, even though both are “2 points.” Treating an ordinal scale like an interval scale to calculate how much someone improved is shaky ground — swap in a different scale, and the conclusion about who improved more can flip entirely.

Checking what AI hands you, or what a study hands you, was never just a matter of whether the conclusion sounds good or how often it’s been cited. It means asking, first, what scale produced that number, and what moment in time it was measuring. Both studies’ “compression” findings deserve a question mark — not because compression doesn’t exist, but because that question matters more than whether compression exists at all.