The companies that learned this the hard way did not have a technology problem. They had failed to ask which decisions a machine should be allowed to make, and which it should hand back.
In late 2023, the Swedish fintech Klarna stopped hiring and began telling the world a clean story about the future of work. Its chief executive announced that artificial intelligence could already do the jobs that humans do, and the company’s figures seemed to prove it: an OpenAI-powered assistant was reported to be handling the workload of seven hundred customer service agents, handling two-thirds of incoming requests, and estimated to drive a forty million dollar profit improvement for 2024. It was the most quotable case of replacement on offer. By May 2025 the same chief executive was telling Bloomberg that the company had gone too far, that the relentless focus on cost had produced lower quality, and that Klarna would be rebuilding its human support. The reversal is the most instructive part of the story, and it is widely misread. Klarna did not discover that AI does not work. It discovered that it had automated the wrong thing.
What had failed was not the technology but the framing. Klarna had treated customer service as a volume of tasks to be removed from human hands, when the work was in fact a distribution of decisions of wildly varying consequence. The routine queries, which an agent handles well, made up a large share of the volume but only a small share of the value at risk. The complex and emotionally charged cases, a smaller share of volume, carried most of the risk to the brand. By automating across that distribution indiscriminately, the company optimized the part that was cheap to get right and degraded the part that was expensive to get wrong. What it eventually settled into, a hybrid in which agents handle high-volume routine requests and humans handle escalations and high-stakes interactions, was not a retreat from AI. It was the decision architecture it should have designed at the outset.
The Decision Is the Unit of Work
The Klarna error is so common because it is built into the language we use. When we frame AI adoption as a question of which tasks can be automated, we inherit a unit of analysis from an earlier industrial logic. Tasks are discrete, repetitive, and bounded; they were the natural thing to optimize when the object was physical throughput. But the work that determines whether a modern organization succeeds is rarely task work. It is decision work: the continuous activity of classifying what has arrived, routing it to the right place, escalating what exceeds normal parameters, approving what meets a threshold, and learning from the accumulated record of all of it. This is the work agentic systems are genuinely suited to, and it is the work conventional automation could never touch, because decision work resists the rigid scripting that automation requires.
The distinction changes what is being optimized. A task-centric view asks how quickly a thing can be done. A decision-centric view asks something harder and more valuable: whether the right thing is being decided, by the right mechanism, with the right escalation when judgment is required, and with a record that lets the organization improve. Agentic AI is the first technology to make the second question operationally tractable at scale. The Gartner research arm has projected that by 2028 agents will make as much as fifteen percent of day-to-day operational decisions autonomously. The question is not whether that share arrives. It is whether the decisions underneath it have been designed, or merely delegated.
Klarna did not discover that AI does not work. It discovered that it had automated the wrong thing.

Five Verbs That Run Every Organization
It is worth being concrete, because abstraction is where these conversations usually go to die. Nearly every operational organization runs on five verbs. It classifies: deciding what kind of thing has arrived, a routine claim or a fraudulent one, a simple inquiry or a regulatory matter. It routes: sending each classified thing to the function equipped to handle it. It escalates: recognizing when something exceeds the routine and requires higher judgment. It approves: applying a threshold and committing the organization to an action. And it learns: feeding outcomes back into the criteria that govern the next round.
These five verbs are where agentic AI delivers value, and the value is not primarily speed. It is consistency in classification that no longer depends on which individual handled the case. It is routing that reflects the actual structure of expertise rather than the accident of who answered first. It is escalation governed by explicit criteria rather than the variable confidence of whoever is on shift. It is approval that applies the same threshold every time, with a record of why. And it is learning captured in the system rather than lost when an experienced employee leaves. None of this removes the human judgment at the center of the hardest decisions. It surrounds that judgment with infrastructure that makes it more reliable and more reviewable. Klarna’s hybrid is precisely this: the agent owns classify, route, and the routine approvals; the human owns the escalations the system is built to recognize and hand off.
Relocation, Not Removal
This is the reframe that matters, and it is why replacement is not merely an inaccurate description but a strategically harmful one. Agentic systems are excellent at the high-volume, well-specified portions of decision work and unreliable precisely where human organizations earn their value: the ambiguous case, the unprecedented situation, the decision that requires weighing considerations no rule anticipated. A well-designed system does not eliminate the human role. It relocates it, concentrating human attention on the exceptions and the genuinely novel while delegating the volume that was previously consuming that judgment in undifferentiated triage. An organization that approaches agentic AI as a headcount exercise will measure success by how many decisions it removed from human hands, and will build a brittle system that handles the routine well and the exceptional catastrophically. An organization that approaches it as decision redesign asks a better set of questions: which decisions to fully delegate, which to have the system make but review, which to escalate by default, and how those boundaries should shift as the system learns.
The Honest Counterargument
We should be fair to the opposing case rather than caricature it, because in its strongest form it is not wrong. Some jobs really do disappear. Klarna’s workforce shrank by roughly forty percent over this period, and the company has been clear that it does not intend to reverse that; its rebuilt human support is smaller and differently shaped than what preceded it. When decision work is redesigned well, the routine layer that once required many people genuinely requires fewer, and it is no comfort to those displaced to be told that the surviving roles are more cognitively demanding. The honest version of our argument is not that employment is unaffected. It is that the value of these systems, and the difference between the organizations that capture it and those that suffer the Klarna reversal, lies in the design of the decision boundaries rather than in the raw subtraction of people. Treating the headcount number as the measure of success is exactly the mistake that produces the brittle outcome. The reduction, where it occurs, should be a consequence of better-designed decision work, not its objective.
Automation does not transfer liability. It concentrates it against the organization that deployed the agent.
The Governance Question Hiding Inside the Technical One
Redesigning decision work surfaces governance questions organizations have historically been able to avoid, because the decisions were distributed across many people and many moments and never had to be specified in one place. The evidence that they are not ready is now quantitative. In one 2025 survey of enterprise IT professionals, eighty-two percent reported using AI agents while only forty-four percent had formal governance policies in place, and eighty percent reported that their agents had already taken unintended actions. Gartner has projected that more than forty percent of agentic AI projects will be canceled by the end of 2027, citing inadequate risk controls among the leading causes. The questions underneath these numbers are not technical. Who is permitted to define the classification criteria an autonomous system applies at scale? What is the escalation threshold, and who owns the consequences when it is set too high or too low? When the system approves something it should have escalated, where does accountability sit, with the engineer who built it, the manager who set the parameters, or the executive who authorized the delegation?
The clean way to state the stakes is that automation does not transfer liability; it concentrates it. The more decisions an agent makes on behalf of an organization, the more those decisions accumulate against the organization that deployed it. This is why the most important work of agentic adoption is frequently the least technical. Before a single agent is deployed, an organization serious about decision redesign will have mapped its decision flows, made its escalation criteria explicit, assigned ownership for each class of decision, and established the mechanism by which the system and its supervisors learn from outcomes. Organizations that skip this work do not avoid it. They encounter it later, after deployment, when an autonomous system has been making consequential decisions according to criteria no one fully specified and no one clearly owns, which is a fair description of how Klarna arrived at its reversal.
What the Best Organizations Will Build
The organizations that extract the most value from agentic AI over the coming decade will not be the ones that automated most aggressively or deployed the most capable models. They will be the ones that used the demands of automation as an occasion to understand their own decision-making for the first time, and to redesign it deliberately. They will treat the five verbs as a system to be engineered rather than a set of habits to be inherited. They will know which of their decisions are routine enough to delegate and which are too consequential to remove from human hands. And they will have built the connective layer, between the decisions a machine can make and the judgment only a person can supply, that turns autonomous capability into reliable institutional performance, rather than a public reversal eighteen months later.
This is harder than buying technology, and slower, and far less amenable to a confident projection of savings. But it is the work that actually matters, because the constraint on the value of agentic AI was never the sophistication of the agent. It was the clarity of the decision work the agent was asked to perform. The companies that learned this the expensive way have left a usable lesson for everyone else: the question worth asking is not how much work a system can absorb, but whether, now that we must finally say out loud how we decide, we can decide better than we ever have before.
SOURCES
- Klarna AI assistant handles two-thirds of customer service chats in its first month; 700-agent workload and estimated $40M 2024 profit improvement. OpenAI / Klarna press release, February 2024. openai.com/index/klarna | prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html
- Klarna CEO Sebastian Siemiatkowski says the AI-first strategy went too far and produced lower-quality service; company rebuilds human support. Bloomberg, May 2025, as reported by Fast Company and CX Network. fastcompany.com/91039401/klarna-ai-virtual-assistant-does-the-work-of-700-humans-after-layoffs
- Klarna workforce reduced by roughly 40 percent over the AI-adoption period. Reuters reporting on Klarna headcount and IPO filings, 2024-2025. reuters.com (search: Klarna workforce AI headcount)
- Production-database deletion by an autonomous coding tool during a code freeze, cited as an agentic accountability failure. Fortune, 2025. fortune.com (search: AI agent deletes production database)
- 82 percent of organizations use AI agents while only 44 percent have policies to secure them; 80 percent report agents taking unintended actions. SailPoint, AI Agents: The New Attack Surface, May 2025 (survey conducted by Dimensional Research). sailpoint.com/press-releases/sailpoint-ai-agent-adoption-report
- More than 40 percent of agentic AI projects projected to be canceled by end of 2027 due to inadequate risk controls; agents projected to make ~15 percent of day-to-day operational decisions by 2028. Gartner, 2025. gartner.com/en/newsroom (June 2025 agentic AI projections)