Why Companies Can't Make AI Value Stick
And how the organisations that succeed are redesigning work instead of automating it.
When companies today struggle with getting value out of their AI investments, it is far from a new problem. In 1987, the economist Robert Solow noticed something strange. America had spent billions on computers. Yet none of it was showing up in economic growth.1 The pattern became known as the productivity paradox.
When factories first electrified in the 1880s, they swapped in electric motors and left everything else the same. For forty years, the productivity data barely moved at all. The value only appeared when a new generation of owners redesigned the factory itself: open floor plans, assembly lines, new roles, new skills. The result was not a more efficient version of the old factory. It was a different kind of factory entirely.2,3
Conceptual model based on David (1990) and Damron (2025)
AI is currently in the electric-motor-swap phase, except the cycle that took electricity forty years is compressing into years, or even months. In early 2026, a survey of nearly 6,000 senior executives across four countries found that over 80% of firms reported no impact of AI on productivity.4 The largest enterprise Copilot evaluation to date (20,000 employees, three months) reported significant user-reported time savings, but the report itself admitted it could not identify how that saved time was actually spent.5
AI does increase productivity in specific cases. Klarna reported that its AI assistant handled 2.3 million customer conversations in a single month, doing the equivalent work of 700 human agents.6 But even that is one task, in one function, designed for automation.
The question is not whether AI can make individuals more efficient. It is why that value evaporates on the way to organisational value.
There is real value in AI-assisted productivity. Nobody disputes that. But there is a limit to what can be achieved by adding AI to existing workflows. Beyond that limit, the work itself has to change.
As with electricity, capturing the value requires building new organisational capabilities. That starts with redesigning work end-to-end: from what AI automates, through the behavioural change each role needs to make, to the new behaviours required for freed capacity to create value. If any link in that chain is undesigned, value leaks out.
The Three Conditions Required for AI Value Realisation
In every previous wave of technology adoption, three behavioural conditions consistently separated organisations that created real value from those that did not.2,7 They are the same conditions that will determine the success of AI programmes.
- The AI must improve the process.
- People must adopt new ways of working.
- Freed capacity must be invested, not just freed.
Conceptual model. Informed by Zuboff (1988) and David (1990)
Are you automating the right work?
AI must improve the process
Before asking whether the AI output is good enough, there is a prior question organisations must ask: does this automation remove a real constraint? Take a strategy consultant who uses AI to produce a 30-page market analysis instead of 10. But the client never read past page 5. The bottleneck was the consultant's ability to distill complexity into a recommendation, not produce more of it. A finance team generates weekly dashboards that used to take a month. But the organisation was already drowning in reports nobody acted on. The constraint was getting a decision-maker to spend 20 minutes with the data and make a call. Faster reports just pile up at the same downstream bottleneck.
And even when AI removes a genuine bottleneck and saves real time, the value is never in the task it performed. It is in what becomes possible because that task is done. A dishwasher's value is not clean dishes. It is the hour the chef got back. But the restaurant only benefits if that hour goes toward a better menu, not more of the same dishes.
Without a named business problem for each team, and visible change downstream of the automation, efficiency stays inside the task. It never becomes organisational value.
Your highest-skilled people are your lowest adopters
People must adopt new ways of working
Here is where it gets difficult. The people best positioned to validate AI output, the senior practitioners, the domain experts, the people with fifteen years of experience, are often the lowest adopters.8 The resistance makes sense. A senior analyst watching AI generate a financial model in seconds is watching the market value of their core skill collapse in real time. Research on social threat explains why AI hits harder than previous technologies:9 it can simultaneously threaten how people feel about their status, certainty, autonomy, relatedness, and fairness. Most previous technologies triggered one of these dimensions. AI can trigger all five at once.
Conceptual model based on Rock (2008)
Most organisations treat the identity threat as a problem to manage. However, Rock's framework suggests the opposite: treat it as intelligence. Help people see where their value is going, not where it is disappearing. That senior analyst's new value is knowing whether the model is right, what it means for the business, and what to do about it. The work does not disappear. It transitions into a higher degree of judgment over what AI produces.
Judgment here is not abstract. It is knowing what the AI was never told: the client relationship, the regulatory shift, the reason last quarter was an anomaly. It is spotting when AI produces a confident answer that misses context only a human would know, the kind of output that would quietly embarrass the business if it reached a client. Judgment is turning a model's answer into a recommendation someone can act on.
AI value depends on the judgment of the people it threatens most.
You cannot make someone a good reviewer by telling them to review. Being excellent at building a financial model does not prepare someone for critically evaluating one. Research on earlier waves of automation found that organisations that invested in building the judgment capability produced more valuable workers. Those that just automated and expected people to figure it out produced worse business outcomes.10
The judgment capability cannot be taught in a workshop. It must be practiced in settings where the new role can be tried before it becomes a requirement.11 Without helping the people whose expertise matters most navigate the shift from producing work to judging it, adoption will be wide but shallow. The tools get used. The behaviour does not change.
The forty minutes nobody can account for
Freed capacity must be invested, not just freed
It's a Tuesday afternoon. Using AI, a marketing analyst finishes a campaign report in twenty minutes instead of the usual hour. Nothing is designed to happen next. So the freed capacity drifts towards email, a few messages, a meeting that could have been a Teams message. A few hours later, the freed forty minutes have vanished.
People are not undisciplined. The system simply removed one prompt for action without replacing it with another. When technology removes a task from someone's daily routine, it removes the prompt for one behaviour without supplying a prompt for the next.
Conceptual model based on Parkinson (1955)
Two factors are responsible for freed capacity disappearing:
First there is Parkinson's Law12: work expands to fill the time available. AI generates a first draft in minutes. Instead of shipping it, the person spends an hour refining because the time exists. The draft was good enough. The extra polish adds nothing the recipient will notice. The same applies in reverse. AI produces an analysis in seconds. Instead of making the decision, the team requests three more scenarios because they're now cheap to produce. The task that was supposed to take an hour still takes an hour.
The second dynamic responsible for freed capacity disappearing is one that Jevons13 helps explain. In the current AI discussion, the Jevons Paradox refers to how cheaper and more efficient AI leads to massively increased AI consumption overall. But inside an organisation, a related effect plays out at the task level: when AI makes something cheap to produce, people produce more of it. AI drafts emails in seconds, so people send more emails, which generates more emails for everyone else to respond to. AI makes it easy to run analyses, so teams produce more analyses without improving a single decision. AI generates presentations faster, so more decks get created, which means more meetings to present them and more people pulled in to review them. The efficiency gain becomes task volume increase. These uncoordinated productivity gains may look like progress individually, but without deliberate design they do not compound into organisational value.
Conceptual model. Inspired by Jevons (1865)
Both consume the freed capacity without producing organisational value. And both are preventable, but only through deliberate design. Without a pre-designed answer to "what do I do now?", people will either produce more of what they already know how to produce, or default into whatever is familiar. How difficult it is to fill that vacuum depends on the distance between the old task and the desired new one. And for any given team, that distance is different for every role.
Without a designed destination for freed capacity and a designed path for each role to get there, the value will vanish into the routines it was meant to replace.
Why the Same AI Deployment Creates Three Different Problems
Understanding the three conditions is the starting point. But applying them requires visibility into where friction will be highest across roles and teams.
A single AI deployment on a single team can create three distinct levels of friction, depending on how far each role must travel from old behaviour to new. The junior developer's transition from routine coding to AI-assisted coding might be low friction. The mid-level developer's transition from writing to reviewing AI output is medium friction, requiring a cognitive architecture they may never have practiced. The senior architect's transition from production to pure judgment and system design is high friction, with real identity-level resistance.
Conceptual model based on Rock (2008) and Ibarra (1999)
But individual friction is only half the picture.
When organisations look at the team as a system, something else becomes visible. Not just where value evaporates, but which new configurations become possible. Maybe junior developers absorb review tasks that senior architects used to handle, tasks they can manage now because AI assists their judgment. This frees the architects for system design work that was never possible before, because they were permanently buried in review queues. The value is not in any individual's freed hours, it is in the new configuration of the team: a structure that produces capabilities the organisation did not have before.
The Measurement Gap
Applying the three conditions also requires the ability to measure whether behavioural change is actually happening. Most AI business cases start from the tool: "Copilot saves developers five hours per week." That is a statement about a technology with no connection to a business outcome.
An honest business case starts from the capability, not the tool. Not "Copilot saves five hours per week" but "We need a 30% reduction in legacy system incidents over twelve months. What does AI need to deliver, what behaviours need to change, and where does freed capacity need to go to make that happen?"
That requires knowing whether behaviours actually changed, not just whether the tool was used. Shallow metrics cannot tell you this. They track activity: logins, clicks, sessions. But the question is behavioural: did someone start their work from the AI draft instead of a blank page? Was freed capacity reinvested in higher-value work or did it dissolve into email? AI interaction data can reveal this at a level that was never possible before. Turning it into a diagnostic requires investment, but the signal has never been more granular.
Conceptual model. Inspired by findings from METR (2025, 2026)
The problem with self-reporting, and other shallow metrics, is not just imprecision. It can be backwards. METR's 2025 trial found developers who believed they were 20% faster with AI were actually 19% slower.14 An era where AI triggers identity-level resistance across entire teams demands deeper behavioural insight than usage dashboards can provide.
What we need to know is not just who is using the tools, but how: which roles have genuinely changed their workflow, who are using AI to do the same work faster, and where freed capacity is actually going. That visibility turns a vague adoption number into a role-level diagnosis that tells you exactly where to intervene and whether the intervention worked.
Six questions that determine whether AI creates organisational value
Most AI programmes find it difficult to answer six behavioural questions about whether their investment is creating real value. Not because the questions are complex, but because the answers require a kind of visibility most organisations have not built: connecting AI interaction data with organisational context to see what is actually changing. We call this visibility the AI Value Baseline.
Together, these steps produce a structured view across the three conditions: which processes genuinely improved, where behavioural change is real versus reported, and whether freed capacity is reaching its intended destination. That view is the starting point for every decision that follows.
The Lesson from ElectricityBuild the muscle now
The productivity paradox that Solow identified was never solved by deploying better technology alone. It was solved by the organisations that did something more difficult: they redesigned work itself. That meant mapping the friction every role would face, building the judgment capabilities people needed, and measuring behavioural change from the work itself, not only from surveys about it. It meant deciding where freed capacity would go before deploying anything.
Conceptual model based on David (1990) and Brynjolfsson (1993)
The three conditions for AI value described in this article are not new. But AI raises the stakes on each one. Processes change faster than organisations can redesign them. The behavioural shift cuts deeper, threatening status, autonomy, and professional identity all at once. The organisations willing to redesign how work is done, invest in the capabilities that new roles demand, and measure whether the change actually holds are not just capturing more from their current AI investment.
They are building the capability to capture more from every AI wave that follows.
References
- 1. Solow, R. (1987). "We'd Better Watch Out." New York Times Book Review, July 12, 1987.
- 2. David, P.A. (1990). "The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox." American Economic Review, 80(2), 355-361.
- 3. Damron, W. (2025). "Gains from Factory Electrification: Evidence from North Carolina, 1905-1926." Explorations in Economic History, 96.
- 4. Yotzov, I., Barrero, J.M., Bloom, N. et al. (2026). "Firm Data on AI." NBER Working Paper No. 34836.
- 5. UK Government Digital Service (2024). Cross-government Microsoft 365 Copilot experiment (20,000 civil servants).
- 6. Klarna (2024). "Klarna AI Assistant Handles Two-Thirds of Customer Service Chats in Its First Month." Press release, February 27, 2024.
- 7. Brynjolfsson, E. (1993). "The Productivity Paradox of Information Technology." Communications of the ACM, 36(12), 66-77.
- 8. Microsoft WorkLab (2024). "What Can Copilot's Earliest Users Teach Us About AI at Work?"
- 9. Rock, D. (2008). "SCARF: A Brain-Based Model for Collaborating With and Influencing Others." NeuroLeadership Journal, 1(1), 44-52.
- 10. Zuboff, S. (1988). In the Age of the Smart Machine: The Future of Work and Power. New York: Basic Books.
- 11. Ibarra, H. (1999). "Provisional Selves: Experimenting with Image and Identity in Professional Adaptation." Administrative Science Quarterly, 44(4), 764-791.
- 12. Parkinson, C.N. (1955). "Parkinson's Law." The Economist, November 19, 1955.
- 13. Jevons, W.S. (1865). The Coal Question: An Inquiry Concerning the Progress of the Nation and the Probable Exhaustion of Our Coal-Mines. London: Macmillan.
- 14. METR (2025). "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." Becker, J., Rush, N., Barnes, E., Rein, D.
- 15. METR (2026). "We Are Changing Our Developer Productivity Experiment Design" (follow-up).