Inflated productivity. What gets reported is the quantity kind

In this article

  1. What the metric decided not to see
  2. The study that counted both things
  3. Four trades where the tail is the trade
  4. When the inflated figure rules the big decisions
  5. What would need measuring and still isn't measured
  6. You might also like

Definitions · References · Further reading · Elsewhere

I do more. I notice it every morning, when I dispatch in an hour what used to take me three. What I don't know is whether I do better, and I suspect the distance between those two questions is exactly what the productivity metric has decided not to look at. The number rises, the "output per hour" figure shoots up with AI, and a number that counts how many things I produce without asking whether any of them was worth it isn't measuring productivity: it's measuring noise with good press.

What the metric decided not to see

Economic productivity is measured, in essence, by dividing. Quantity produced over hours worked, or quantity produced over number of workers. The OECD, Eurostat and the US Bureau of Labor Statistics have for decades published long series with those two yardsticks, which are the ones we have and are yardsticks of quantity.

For a long time they sufficed. In an economy of steel and highway, quantity correlated reasonably well with value: more tonnes were more bridges, more kilometers were more trade, and one report more cost nobody anything because almost nobody produced reports. That silent correlation between quantity and value is the beam holding up the entire accounting of productivity. It's also the one the cognitive economy has been breaking in half for years without the metric noticing.

Because the yardstick doesn't weigh quality. A mediocre report and an excellent report enter the statistic as two reports, identical. Nor does it weigh need: ten documents nobody asked for count ten; one indispensable one counts one. And since it ignores the distribution, when the average rises while the tail —the critical cases, the rare ones— degrades, the aggregate declares victory precisely where the most is being lost. What it doesn't see, above all, are the costs the fast output pushes out of frame. If what I produce quickly forces whoever receives it to filter the garbage before using it, the cost hasn't evaporated: it's moved to the office next door, where no statistic is going to look for it.

The yardstick is still there, intact, for a reason that's nothing noble. It's what we know how to measure.

The study that counted both things

Before going on it's worth looking at the data, because there's data and not all of it says the same thing, and the difference between them is the whole article.

Brynjolfsson, Li and Raymond observed what happened when a conversational assistant entered a customer support center with more than five thousand agents. The average productivity increase came to around 14%, with jumps of 34% among the least experienced workers and a practically nil effect among the most expert. AI, they argued, seems to diffuse the best practices of the best downward. So far, the story everyone repeats in the presentations.

Dell'Acqua and his coauthors did something rather more uncomfortable. Instead of measuring an average, they separated the tasks by whether they fell inside or outside what the model masters. With Boston Consulting Group consultants and GPT-4, those who used AI on tasks inside that capability frontier were around 25% faster and delivered higher-quality results. But on a problem designed on purpose to fall outside the frontier, in what the model doesn't handle well, the AI users got it right less often than the group working without it. The tool doesn't warn you where the edge is. It pushes with the same confidence to one side and the other.

The pattern reappears in code, and reappears split in two. Peng and his colleagues timed developers writing a server with GitHub Copilot and saw them finish almost 56% faster than the control group. That speed is real and measured with laboratory cleanliness. What the experiment couldn't measure was what happens to that code three months later. That's where GitClear comes in, which instead of timing an isolated task set about reading the history of hundreds of millions of real lines. Its portrait is the exact reverse of the optimism. The churn —lines reverted within two weeks of being written— has been rising steadily; the duplication of code blocks has shot up; and refactoring has sunk below plain copy-paste for the first time. The speed rises. The health of the codebase falls, and it falls precisely in the dimension no productivity stopwatch registers.

There's the asymmetry that matters. It isn't that AI produces a worse average output. It's that it improves where the average rules and washes its hands of —when it doesn't outright degrade— where the tail rules.

Four trades where the tail is the trade

There are jobs whose distribution of criticality is tipped toward the few rare cases. The bulk is routine, and the whole value of the trade lives in the difficult minority. Applying a quantity metric to those jobs isn't being imprecise: it's looking at the opposite side of where the important thing happens.

Programming is the textbook case. Most code is plumbing any model writes without breaking a sweat, and then there's that fraction where the expensive bugs live, the architecture decisions that hold up the whole system, the security component that can't fail because if it fails it isn't measured in minutes but in breaches. AI accelerates the plumbing, yes. On the rest, Apiiro's research into assisted code points to a rise in security findings and attack surface as the volume generated grows. The productivity statistic sees the faster plumbing and doesn't see the breach, because the breach isn't an output: it's the absence of one that should have existed and didn't.

The medical diagnosis has the same silhouette. The general consultation dispatches the typical in a chain, and the doctor's judgment is truly earned in the infrequent cancer, in the atypical presentation of a common disease, in the comorbidity that fits no protocol. I don't have a closed figure quantifying how much the attention of a clinician who leans on the model for the rare cases relaxes, and I prefer not to invent it. But the logic of the system doesn't need a figure to be seen: an assistant that gets the typical right ninety percent of the time trains the doctor to trust, and trust is exactly what's in surplus in the atypical case, where what's needed is to distrust.

The law repeats the scheme with clauses. The standard ones are pure repetition and automation dispatches them effortlessly; the distinctive ones are where the negotiation that justifies the fees lives. Scientific review takes it to the extreme: summarizing a paper is trivial and the models summarize beautifully, but detecting the hidden methodological flaw that invalidates the conclusion is the hard part of the trade, and it's precisely what they do worst.

In the four cases the measured productivity rises, because it counts outputs without weighting which of them mattered. The quality in the cases that truly decide the trade falls. And the distance between what the number says and what happens on the operating table, in the contract or in the repository stops being a reasonable margin of error and becomes structure.

When the inflated figure rules the big decisions

What comes now is my reading, not the result of any study, and it's worth saying so before going on. What does have a name is the underlying suspicion, which comes from far back.

Charles Goodhart stated in 1975, regarding British monetary management, an idea that would later become famous: any observed statistical regularity tends to collapse once it's pressed for purposes of control. The catchy formulation almost everyone cites —"when a measure becomes a target, it stops being a good measure"— isn't Goodhart's, but that of the anthropologist Marilyn Strathern, who coined it in 1997 condensing that intuition. The mechanism is always the same: people optimize the measure, not the phenomenon the measure was meant to capture. Move it to AI-assisted productivity. Professionals and companies optimize the measured output, not the real quality, and the indicator turns into an ever-weaker signal about what it claimed to represent.

And here's where an overly optimistic figure stops being a technical detail. Central banks use productivity estimates to infer how much an economy can grow without overheating; if real productivity is lower than reported because quality falls in the tail, interest rates get set on an optimism nobody has verified. The big public investment programs in digitalization —the European Next Generation, the US IRA— finance the adoption of AI counting returns based on measured productivity, so that, if the measure is inflated, the social return they promise is lower than the brochure's. Labor policy calibrates the pace of reskilling and training on sectoral projections that, overestimated, arrive late or are in surplus where they weren't needed. And within each company, headcount, training and investment get decided on an output that lies precisely about the critical cases, which is where the losses nobody is counting concentrate.

It isn't that data is missing. It's that the surplus data has become the one that rules.

What would need measuring and still isn't measured

The yardsticks capable of capturing what the classic ones leave out exist, but in an embryonic state, scattered, not aggregated into any macro dashboard.

In code there's talk of the defect escape rate, the proportion of defects that reach production without having been detected earlier; in medicine, of the adverse clinical outcomes that appear in the long tail; in law, of the rate of clauses that end up generating litigation. They're metrics of sustained quality, not of speed, and they look where quantity doesn't. The problem is that they stay in the specific repository, hospital or firm without ever climbing into an aggregate. Harder still is measuring the cost the fast output loads onto whoever receives it —how much time it costs someone to process what you generate in a minute—, a figure almost nobody reports because it amounts to admitting that one person's speed is another's slowness. And the most demanding of all would be to weight each output by the importance of the case, which forces you to classify the cases by criticality first, a task the system, comfortable in its aggregate, has not the slightest interest in doing.

While those yardsticks don't mature, the productivity that gets reported will keep being the inflated one, the decisions will keep being taken on it, and the gap between the measured and the real will keep widening with the calm of someone reviewing a spreadsheet where everything adds up. It adds up because it only counts what it knows how to count. And what it doesn't know how to count is, more and more, the only thing that mattered.

Definitions

Output per hour worked. Quantity produced divided by hours spent. It's the most-used labor productivity metric in the official series; it measures volume, not value or quality.

Model capability frontier (in English, jagged frontier). The irregular set of tasks where an AI model operates with quality, surrounded by apparently similar tasks at which it fails. A concept operationalized by Dell'Acqua and coauthors (2023).

Long tail of criticality. The minority of cases whose importance is disproportionately high relative to their frequency. A common pattern in professional trades, where the bulk is routine and the value concentrates in the infrequent.

Code churn. Lines of code reverted or rewritten in the weeks following their creation. A high churn suggests code that doesn't hold up and must be redone soon.

Defect escape rate. The proportion of defects that reach production without having been detected earlier. A metric of sustained quality in software engineering.

Goodhart's law. The idea that a statistical regularity stops being reliable once it's used as a control target. Stated by Charles Goodhart in 1975; the popular formulation is Marilyn Strathern's (1997).

References

Brynjolfsson, E., Li, D. & Raymond, L.Generative AI at Work, NBER Working Paper 31161 (2023). A study of more than five thousand customer support agents: average productivity increase close to 14% and of 34% among the least experienced, with a minimal effect among the most expert. https://www.nber.org/papers/w31161

Dell'Acqua, F. et al.Navigating the Jagged Technological Frontier. Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, Harvard Business School Working Paper 24-013 (2023). BCG consultants with GPT-4: around 25% faster and higher quality inside the frontier, worse results on a task designed outside it. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700

Peng, S. et al.The Impact of AI on Developer Productivity. Evidence from GitHub Copilot, arXiv:2302.06590 (2023). A controlled experiment: the group with Copilot finished 55.8% faster. No measurement of long-term code quality. https://arxiv.org/abs/2302.06590

GitClearCoding on Copilot (2024) and AI Copilot Code Quality (2025). Analysis of hundreds of millions of real lines: sustained increase in churn relative to the pre-AI base, increase in code duplication and decline in refactoring below the level of copy-paste. https://www.gitclear.com/ai_assistant_code_quality_2025_research

Apiiro4x Velocity, 10x Vulnerabilities. AI Coding Assistants Are Shipping More Risks (2025). Associates the increase in AI-generated code with a growth in security findings and attack surface. https://apiiro.com/blog/4x-velocity-10x-vulnerabilities-ai-coding-assistants-are-shipping-more-risks/

Goodhart, C.Problems of Monetary Management. The U.K. Experience (1975). The origin of the idea known as Goodhart's law.

Strathern, M.«Improving Ratings». Audit in the British University System, European Review 5 (1997). The source of the popular formulation "when a measure becomes a target, it stops being a good measure."

OECD — Labor productivity series, Compendium of Productivity Indicators. https://www.oecd.org/sdd/productivity-stats/

Bureau of Labor Statistics (US) and Eurostat — Official series of sectoral productivity.

Further reading

Coyle, D.GDP. A Brief But Affectionate History (Princeton University Press, 2014). A critical history of the indicator.

Stiglitz, J., Sen, A. & Fitoussi, J.-P.Mismeasuring Our Lives. Why GDP Doesn't Add Up (The New Press, 2010).

Brynjolfsson, E. & McAfee, A.The Second Machine Age (W. W. Norton, 2014). A prior framework, comparative reading with the later literature.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment