There is a price for a unit of thinking, and it is falling faster than almost any price in the modern economic record. In November 2022, querying a machine that could perform at roughly the level of a competent generalist cost about twenty dollars for a million tokens. By October 2024 the same capability cost seven cents. A two-hundred-and-eighty-fold collapse in eighteen months, and the curve has not flattened since.


Prices do not usually behave this way. The cost of a loaf of bread, a kilowatt-hour, an hour of plumbing, they drift, they do not fall through the floor. When a price detaches from gravity like this, it is worth asking what, exactly, has become cheap. The answer is uncomfortable, because the thing that has become cheap is a passable imitation of the one commodity the developed world had organized its entire labor market around selling: cognitive effort. Reading, drafting, summarizing, answering, reconciling, coding the routine parts. The work of the desk.


And once a thing becomes cheap, capital goes looking for the places it is still expensive.


That is the whole story, mechanically. Somewhere between the salary line on a payroll and the metered line on a cloud-services invoice, a gap has opened. It is a wide, structural, exploitable gap between the price of human cognition and the price of its synthetic substitute. Markets do not leave gaps like that alone. They build machinery to stand inside them and collect the difference. Call it the labor arbitrage machine: not a single tool or company, but the emergent apparatus; models, agents, vendors, procurement teams, investor expectations. All that converts the price gap between a human hour and a thousand tokens into margin.


Engineers have a phrase for what happens to a well-designed system when the load exceeds what it can serve: graceful degradation. The system does not crash. It sheds the least essential functions, runs slower, drops quality at the edges, and stays standing. The labor market is now being asked to degrade gracefully, to shed the most automatable cognitive tasks, quietly, function by function, without anyone declaring a crisis. Totale, because the shedding is not confined to one trade. It reaches the call center and the back office and the junior desk and the first draft, all at once. This essay is about that machine: how it prices its inputs, where it is already running, and the three postures a firm can take toward it – refusal, the loop, and reinvention – each with a real cost and a real casualty.


The Price of a Thought
Begin with the input that makes the machine possible, because without the price collapse none of the rest is even arithmetic. Stanford’s AI Index, drawing on Epoch AI, tracks the cost of running a model at a fixed level of capability over time. The headline series is the one above: from twenty dollars per million tokens to seven cents in a year and a half. Depending on the task, inference prices have been falling somewhere between nine and nine hundred times per year, with the fastest declines arriving after January 2024. The floor keeps dropping as competition thickens. This is a price war between a handful of frontier labs and a long tail of open-weight challengers willing to undercut on raw cost.

Figure 1. The price of a thought, collapsing: the cost to query a model at GPT-3.5 capability fell roughly 280-fold in eighteen months. Source: Stanford AI Index 2025, citing Epoch AI.


What that line means in practice is that intelligence, or at least the part of it that can be reduced to pattern, retrieval, and fluent paraphrase, has been transformed from a scarce, expensive, human-embodied resource into something closer to a utility. Metered. Abundant. Billed by the unit. And like every utility before it, the moment it became cheap and reliable, the question stopped being whether to use it and became where to point it.


Here the two ledgers diverge in a way that organizes everything that follows. A human worker is a fixed cost with a floor: a salary, benefits, a desk, a manager’s attention, sick leave, the slow expensive business of being a person. A model is a variable cost with no floor; you pay per token, the unit price falls every quarter, and when the work stops the meter stops with it. One of these lines bends downward on its own. The other does not. A finance function that has internalized that asymmetry does not need to be told what to do; the spreadsheet tells it.


The Arithmetic of Substitution
Abstractions persuade no one. The machine becomes legible only when you put a single unit of work on the table and price it both ways. Take the most automatable cognitive job in the economy by sheer headcount: the customer service representative. In May 2024 the United States employed roughly 2.8 million of them, at a median wage of about $42,830 a year. The Bureau of Labor Statistics already projects the occupation to shrink five percent over the decade to 2034, a loss of some 154,000 posts, which is itself a quiet forecast about who wins this comparison.


Run the numbers per resolved conversation. A human agent handling perhaps fifty chats a day across roughly two hundred and fifty working days resolves on the order of 12,500 conversations a year. Against base wages alone, that is about $3.43 per conversation, and the true figure is higher once benefits, supervision, premises, and attrition are loaded in. Now price the same conversation as tokens. Assume a realistic agentic exchange: a system prompt, retrieved account and policy documents, a few turns of back-and-forth, call it fifteen thousand tokens in and two thousand out. At early-2026 list prices, that single resolved conversation costs roughly twelve cents on a frontier model, six cents on a capable mid-tier model, and about a fifth of a cent on a budget model.

Figure 2. The arithmetic of substitution. Illustrative cost to handle one resolved customer conversation: human base wage versus per-token model cost at three price tiers. Assumes ~15,000 input + 2,000 output tokens per resolved conversation; excludes build, integration, and human-oversight costs. Sources: BLS (wages); published API list prices, early 2026.


Three cents. Six cents. Twelve cents. Against three dollars and forty-three.


Hold the obvious objections, because they matter and we will return to every one of them: the human handles the furious customer, the ambiguous edge case, the conversation that needs a person to simply be a person; the model needs building, integrating, supervising, and correcting; the cheapest model is not the most reliable; and the comparison flatters the machine by costing only the easy tier of work. All true. But none of it closes a gap of this magnitude. When the per-unit cost of a substitute is one-thirtieth to one-fifteen-hundredth of the incumbent, the substitute does not have to be as good. It only has to be good enough on enough of the volume, while the arbitrage takes care of the rest.


There is a subtler effect, easy to miss if you think only in terms of substitution. When the marginal cost of a cognitive task falls toward zero, it does not merely make the existing task cheaper; it makes entirely new things economic that were not before. Support a firm could never have afforded to staff becomes a free feature. A product that once justified a single tier of service can suddenly justify infinite, instant, multilingual service. Prompt caching and batch processing push the per-unit cost down by a further order of magnitude on repeated work. The arbitrage, then, is not a one-time saving banked against a fixed list of tasks. It is a moving floor that keeps descending.


If that sounds theoretical, it has already been run at scale, in public, by a company willing to put a number on it. In February 2024 the fintech Klarna launched an OpenAI-powered customer-service assistant. Within its first month it had handled 2.3 million conversations, about two-thirds of the company’s total chat volume, doing what Klarna described as the work of 700 full-time agents. Resolution time fell from eleven minutes to under two; repeat inquiries dropped a quarter; the assistant operated in more than thirty-five languages, around the clock. Klarna put the profit improvement at roughly forty million dollars in the first year.


Notice what that figure actually is. Seven hundred agents at a loaded cost of fifty to fifty-five thousand dollars each is, near enough, forty million dollars. Klarna’s headline savings number is the labor arbitrage, stated in dollars, by the firm capturing it. The market understood the signal immediately: on the day of the announcement, shares of Teleperformance, one of the world’s largest call-centre operators, fell sharply as investors repriced the future of an entire industry built on human conversational labor. The machine had been demonstrated, and the demonstration was the point.


Adoption Without Transformation
One demonstration does not make a regime. So how widely is the machine actually running? McKinsey’s 2025 survey of nearly two thousand organizations found that 88 percent now use AI in at least one business function. That is an increase of 78 percent a year earlier and 55 percent the year before that. Adoption, in other words, is no longer a differentiator; it is the baseline. Stanford’s index tells the same story from a different sample: organizational use jumping to 78 percent in 2024 from 55 percent, and the use of generative AI in at least one function more than doubling to 71 percent.


And yet. Beneath the adoption curve sits an uncomfortable second number. Only around 39 percent of McKinsey’s respondents reported any measurable impact on enterprise earnings, and only about six percent qualified as genuine high performers seeing material EBIT effects. The majority sit in what the report’s readers have taken to calling pilot purgatory: experiments that never graduate to production. Nearly everyone has the machine plugged in. Almost no one has rewired the building around it.

Figure 3. Adoption without transformation. AI use is nearly universal and still climbing, but enterprise-level impact remains concentrated in a small minority of firms. Source: McKinsey, The State of AI (2023–2025).


This gap is not a footnote; it is the texture of the moment. The discourse speaks of 2025 as the year of the AI agent; systems that do not merely answer a prompt but plan, decide, and execute multi-step workflows. In the consultancy’s phrase they are known as “digital employees”. Twenty-three percent of firms report scaling such agents somewhere in the organization. But in any given function, fewer than one in ten has agents in genuine production. The machine’s reach, for now, exceeds its grip. Which means the interesting question is not whether firms will pursue the arbitrage but how. And here the field narrows to three postures, three answers to the same spreadsheet.


The workforce expectations buried in the same survey are where adoption and arbitrage finally touch. A median of thirty percent of respondents now expect headcount to fall in the functions where AI is deployed and roughly a third anticipate an enterprise-wide reduction of three percent or more. These are not yet cuts; they are intentions, the leading edge of the spreadsheet becoming policy. The reason most firms have not yet realized the savings is mundane and revealing: they are good at running AI projects and bad at rebuilding the operating model around them. Capturing the arbitrage at scale means redesigning the workflow end to end, not bolting a model onto the old one. The gap in Figure 3 is the distance between buying the machine and rewiring the building.
Three Postures Toward the Machine


A firm staring at the gap between $3.43 and three cents has, in the end, only three things it can do. It can refuse the trade and keep its people. It can split the difference and put a human in the loop. Or it can take the trade in full, shed the labor, and redeploy the freed capital. Each is defensible. Each is being chosen, right now, by serious companies. And each extracts a price that is paid by someone.


I. The Refusal
The first posture is to decline. To treat human judgement, presence, and accountability as the product rather than a cost to be optimized away, and to say so. This is not nostalgia; in several markets it is a strategy. There are domains where the human is the value proposition: private wealth, complex care, high-trust B2B relationships, the bespoke and the consequential. A bank that automates the furious client at the worst moment of their financial life has not saved money; it has taught that client to leave. The refusal banks on the premise that as synthetic interaction becomes infinitely abundant, a real human becomes a luxury good.


The history of automation offers more examples of this dynamic than either its advocates or critics usually admit. Mechanical looms did not eliminate handmade textiles; they transformed them into luxury goods. Industrial agriculture did not eliminate artisanal food production; it narrowed it into premium markets whose value derived precisely from their departure from industrial scale. Recorded music did not eliminate live performance. In each case abundance increased the relative value of scarcity. The mass-produced product won the volume market, but the handcrafted product often captured the margin.


The same possibility exists for cognition.


As synthetic intelligence becomes abundant, instantaneous, and nearly free, certain forms of human judgement may acquire value precisely because they remain expensive. A legal opinion signed by a named partner carries a form of accountability that no model can assume. A physician delivering a difficult diagnosis provides not merely information but responsibility for that information. A private banker managing generational wealth sells trust as much as analysis. In each case the customer is purchasing not only the answer but the person attached to it.


There is evidence that this distinction already matters. Studies of automation in medicine repeatedly find that patients express lower trust in fully automated decision-making, particularly in consequential domains involving diagnosis, treatment, or uncertainty. Trust rises when AI serves as an advisor rather than a replacement, preserving visible human accountability. Likewise, surveys of financial and legal services continue to show that clients place disproportionate value on access to identifiable experts even when routine analysis becomes commoditized. The technical quality of the answer matters; the ability to assign responsibility for it matters more.


This creates an unusual economic possibility. The firm that refuses full automation is not necessarily rejecting efficiency. It may instead be repositioning itself into a different market entirely. When machine-generated interaction becomes ubiquitous, human interaction becomes differentiated. The customer who cannot tell whether they are speaking to a machine is unlikely to pay a premium. The customer who knows they are speaking to a person sometimes will.


The refusal therefore treats humanity not as a cost centre but as a scarcity asset. Its wager is that while cognition can be automated, accountability cannot; while information can be reproduced infinitely, trust remains stubbornly attached to individuals. The firm choosing refusal is not betting against technology. It is betting that abundance creates its own demand for authenticity.


The supporting evidence is not merely sentimental. Recall that the arbitrage flatters the machine by costing only the easy tier of work; the moment volume tilts toward the ambiguous, the empathetic, or the legally consequential, the model’s failure rate becomes the firm’s liability. Tacit knowledge – the practical understanding learned on the job, never written down, and therefore absent from the text the models were trained on – remains stubbornly human. There is a reason the senior practitioner is so hard to automate: the thing she knows was never typed anywhere for a model to read.


The asymmetry of failure is the refusal’s strongest card. A human agent who mishandles a conversation does so once, quietly, to one customer. A model deployed across the whole volume can fail in the same way to thousands at once. In the age of the screenshot, a single grotesque failure travels further than a year of competent service. The firm that automates its most sensitive interactions is not trading a small per-conversation saving against a small per-conversation risk; it is concentrating a diffuse, survivable human error rate into a single systemic one. For a brand whose entire value rests on trust, that can prove catastrophic at the precise moment it looks most efficient.


But refusal has a cost, and it is brutal in its simplicity. A competitor who takes the trade can underprice you on the entire automatable tier of the market and use the margin to compete for the rest. Hold the line on human-only service and you may find yourself defending a premium that the market has quietly decided not to pay. As I have written before in another context, a neo-Luddite in 2026 only has his future to miss while the world accelerates around him. The refusal is honorable; it is also, for any firm exposed to price-sensitive volume, a wager that the gap will not widen. The gap is widening.


II. The Loop
The second posture refuses the binary. It keeps the human and the machine in the same workflow – the human in the loop – and tries to capture most of the efficiency without surrendering judgement, accountability, or the apprenticeship that produces the next generation of competence. This is the posture the evidence treats most kindly, and it is the one Klarna itself eventually adopted after overreaching.


The foundational study here is Brynjolfsson, Li and Raymond’s Generative AI at Work, which tracked the staggered rollout of a conversational AI assistant across 5,179 customer-support agents handling some three million chats. Access to the tool raised productivity, defined as “issues resolved per hour”, by 14 percent on average. But the average concealed the finding that matters: a 34 percent improvement for novice and low-skilled workers, and close to zero for the experienced and highly skilled. The model worked by capturing the tacit know-how of the best agents and disseminating it to the rest, helping newer workers move down the experience curve faster. It also improved customer sentiment and raised employee retention.


One detail from that study deserves to be set apart, because it is the entire argument for the loop in a single number. The agents adopted only about 38 percent of the model’s suggestions. They were not stenographers transcribing an oracle. They exercised judgement over which advice was good. The system’s value came precisely from the combination of the model’s breadth and the human’s discrimination. The loop is not a compromise between two inferior options. On the evidence, it can be better than either alone.


The machine raises the floor. The human still has to choose.


Klarna’s second chapter is the cautionary half of the same lesson. Having declared the work of 700 agents automated, the company cut its workforce sharply from around five thousand to roughly thirty-five hundred, largely through attrition. Then, in May 2025, its chief executive told Bloomberg the company had cut too deep on humans, and began reopening hiring for premium human support. The most-cited automation success story in the industry had quietly rebalanced toward the loop. The lesson it drew in public was the one the data already implied: automate the high-volume, low-empathy tier, and move the humans up the value chain rather than out of the building.


But the loop carries its own concealed cost, and it is the one I find most worth dwelling on, because it does not show up on any spreadsheet for years. If the model most helps the novice, lets the beginner perform like a veteran on day one, then the firm’s incentive to actually train that beginner quietly evaporates. Why invest in the slow, expensive manufacture of competence when a subscription delivers a passable version of it instantly? This is the skills-ceiling problem, and it connects directly to something I have argued at length elsewhere: the apprenticeship layer of cognitive work – the beginner tasks through which novices have always become professionals – is precisely the layer the machine performs most fluently. Raise the floor for the novice today and you may find you have removed the staircase by which novices once became experts. The loop preserves the human in the room. It does not, on its own, preserve the path that put her there.


There is a second cost to the loop, quieter than the first. A human-in-the-loop system is, by construction, a measured human. Every suggestion accepted or refused, every resolution time, every sentiment score becomes data. The same apparatus that augments the worker also surveils her at a granularity no clipboard ever achieved. I have written before about how digital systems reduce human interaction to quantifiable engagement; the loop imports that logic into the workplace itself, turning judgement into a metric and the worker into a node whose deviations from the model can be flagged. The augmented worker is more productive. They are also more legible, more comparable, and should the arbitrage ever tilt far enough, more easily dispensed with. Augmentation and replacement are not opposites; often they are the same project at different stages.


III. The Reinvention
The third posture takes the trade in full and is unapologetic about it. It treats AI not as an assistant to existing workers but as a reason to have fewer of them. It reframes the resulting layoffs not as loss but as reinvention: capital freed from routine cognitive labor and redeployed toward higher-value work. This is the posture that generates the headlines, and the fear.


The displacement is real and it is being named. The outplacement firm Challenger, Gray & Christmas, which has tracked AI as a stated reason for job cuts since 2023, recorded 54,836 layoffs explicitly attributed to AI in 2025 alone. That is more than triple the combined total of the two prior years, and part of a cumulative figure of nearly seventy-two thousand since tracking began. Against a year in which US employers announced roughly 1.17 million cuts in total, the highest since the pandemic, and in which the technology sector alone shed over 154,000 posts, the AI-attributed number looks almost modest. Almost.

Figure 4. The named cause is rare and rising. US layoffs explicitly attributed to AI more than tripled from the 2023–2024 combined total to 2025, and are widely thought to undercount the real effect. Source: Challenger, Gray & Christmas.


But the named figure is the wrong number to watch, and the firm that compiles it said so plainly. “Regardless of whether individual jobs are being replaced by AI,” its workplace expert observed in early 2026, “the money for those roles is.” That sentence is the reinvention posture in miniature. Firms rarely announce that a chatbot fired a department. They announce restructuring, efficiency, a pivot toward AI investment. The budget that used to fund a row of desks is rerouted to compute and to a smaller number of higher-paid roles. The arbitrage rarely appears in the layoff notice. It appears in the capital-expenditure line.


IBM is the instructive case, because it shows the reinvention working exactly as advertised and reveals what that means. The company used AI agents to automate the bulk of its routine human-resources tasks; its internal system, by its own account, now handles the overwhelming majority of such requests, and a few hundred HR roles were displaced. Yet IBM’s total headcount rose. “Our total employment has actually gone up,” its chief executive told the Wall Street Journal, “because what it does is it gives you more investment to put into other areas” such as programmers, salespeople, the “critical-thinking” roles where humans face other humans rather than process rote work. This is the optimistic ledger: automation as a capital-reallocation engine, shrinking the routine and funding the strategic.


The optimistic case has history on its side, and it deserves to be stated at full strength. Every previous wave of automation provoked confident predictions of permanent unemployment, and every previous wave was falsified by adaptation: new industries, new roles, work no one had thought to want. The reinventionist bets that this time is no different; that the surplus released by cheap cognition funds the next thing, as it always has. The wager is reasonable. But it rests on a structural assumption this particular wave quietly violates. For two centuries, automation climbed the ladder from the bottom rung of physical drudgery upward, leaving cognitive work for last. Generative AI has inverted the climb: it is strongest at precisely the entry-level cognitive tasks through which beginners once became professionals. Adaptation needs a pathway, and a pathway needs beginners permitted to exist long enough to turn into something else.


And it is genuinely optimistic for IBM, and for the programmers it hired. The harder questions are distributional, and they are the ones the reinvention posture is least eager to answer. The HR clerk whose work was automated is not, in general, the person hired into the better-paid engineering role. The surplus the arbitrage releases is real; the claim that it accrues to the displaced is mostly aspirational. Which returns us to a point I have made before about a different transfer of value: in a system under structural pressure, value does not disappear – it is captured. The arbitrage generates a surplus. The only question that matters is who captures it.


A price advantage, however dramatic, is not a law of nature. Economic history contains numerous cases in which a technically superior or cheaper alternative spread far more slowly than its arithmetic suggested. Nuclear power promised electricity too cheap to meter, yet regulation and public opposition constrained its deployment for decades. Electronic health records took years to penetrate healthcare despite obvious efficiency gains because institutions, incentives, and workflows resisted change. Even the shipping container, perhaps the most economically transformative logistics innovation of the twentieth century, spent years waiting for ports, unions, railroads, and regulators to reorganize themselves around it.

The lesson is not that economics loses. It is that economics often travels at the speed of institutions. A sufficiently large price gap exerts relentless pressure toward adoption, but pressure and transformation are not the same thing. The machine may be economically inevitable without being organizationally immediate.


The Terms of the Trade
If the arbitrage is going to run regardless, which the price gap guarantees it will, then the useful work is not to denounce it or to cheer it, but to argue over its terms. The machine is indifferent to how its surplus is distributed. People are not, and neither, in the end, are the economies they compose. A few principles follow from everything above, addressed to the parties who can actually set those terms.


For firms: the loop beats the binary on the evidence we have. The reflex to chase the cheapest model into the largest layoff ignores both the failure-rate liability and the slower erosion of the talent pipeline. The serious move is to automate the genuinely routine, route the consequential to humans, and keep manufacturing competence on purpose, because the firm that stops training beginners is borrowing its senior talent from a future that may not arrive. The World Economic Forum’s employers already report that 39 percent of workers’ core skills will be outdated between 2025 and 2030, and name the skills gap as their single greatest barrier to transformation. Reskilling is not corporate charity. It is the maintenance schedule for the only input the machine cannot yet supply.


For workers: the defensible ground is the work the model performs worst: judgement, accountability, the management of other humans, and the tacit craft that was never written down. The IMF estimates that around 40 percent of jobs globally are exposed to AI, rising to roughly 60 percent in advanced economies, where about half of the exposed roles may be harmed rather than helped. Exposure is not destiny, but it is a map. The instinct to compete with the machine on speed and volume is a losing one; the instinct to climb toward what it cannot do is not.


For policymakers: the apprenticeship problem is a public problem, because no individual firm has an incentive to solve it. For most of economic history the transition from novice to professional was engineered, not left to chance. Guilds, indentures, the post-war ladder of cheap training and internal promotion; all those provided a pathway to professional augmentation and amelioration. If the machine is dismantling the entry-level rung faster than the market is rebuilding it, then some modern equivalent of the indenture, like subsidized training posts, public apprenticeships, a deliberate manufacture of the next cohort, is not nostalgia but infrastructure. Transparency helps too: if firms were required to name AI as a cause of displacement rather than laundering it through “restructuring,” we would at least be arguing over a real number instead of an undercount.


And the distribution question is not merely one of fairness; it is one of stability, which even the most hard-nosed reinventionist has reason to weigh. A cohort that does everything it was told; accumulates the credentials, masters the tools and still finds the entry-level rung sawn off, does not stay quiescent forever. Thwarted expectation is among the most reliable raw materials of unrest. An economy that captures the arbitrage narrowly, routing its surplus to capital and a thin tier of senior labor while hollowing out the path beneath them, is not only being unjust. It is manufacturing a frustrated electorate, and billing the invoice to a later decade.


The arbitrage is a fact. Its terms are a choice.


Degradation, Gracefully
Return to the word. Engineers prize graceful degradation because the alternative – the sudden, total, catastrophic failure – is so much worse. A system that sheds load quietly and keeps standing is, by the standards of engineering, a success. That is precisely what makes the economic version so difficult to resist, and so easy to wave through. There is no crash to point at. No date on which the labor market fell over. Only a long, soft shedding: the call centre thinned, the back office automated, the junior desk never filled, the first draft now machine-made. Each decision is locally rational, priced by the same widening gap, however none of them announcing themselves as a turning point.


The machine does not hate the workers it displaces, any more than a spreadsheet hates the cell it overwrites. It simply notices a price difference and collects it. That is its nature, and the price difference is not going to close; the line in Figure 1 does not bend back up. What remains undecided is not whether the arbitrage runs but who writes its terms. That is, whether the surplus it releases is captured narrowly or shared, whether the staircase from novice to expert is dismantled or rebuilt, whether the human is moved up the value chain or simply out of the building.


Degradation totale, then, is not a prophecy of collapse. It is a description of a posture: an economy quietly optimizing away the most automatable fraction of human cognitive work, function by function, gracefully, without ever quite admitting that is what it is doing. The system stays standing. The feeds refresh. The margins improve. And somewhere in the slow shedding, a category of human effort is being repriced toward zero; not with a crash, but with a discount.


Keep cool. Watch the price of a thought. The machine is already running.

Posted in

Discover more from Perchance to Economics

Subscribe now to keep reading and get access to the full archive.

Continue reading