Companies Rein In AI Use as Compute Costs Surge

The Financial Times documents a quiet pullback at Amazon, Walmart, Cisco, Uber and Meta. The same story turns every prompt into a metered data point.

A calculator resting on top of printed financial documents on a wooden desk, the kind of paperwork where enterprise AI cost decisions are being finalized in the second quarter of 2026
Photo via Unsplash

TL;DR: Companies that raced to put generative AI in the hands of every worker are quietly pulling back. Amazon, Walmart, Cisco, Uber and Meta are capping internal AI tool budgets, pushing employees to cheaper models, and warning staff against "AI for the sake of AI." The Financial Times reported the pullback on 19 June 2026, citing named executives at each company and a CFO-level shift in how compute cost is being tracked.[1] Uber blew through its entire 2026 AI budget by April and now limits employees to $1,500 per month in token spend on individual tools. Walmart has capped tokens on its internal "Code Puppy" vibe-coding platform after usage "really skyrocketed." Amazon warned staff last month to stop using AI "just for the sake of using AI" after engineers deployed agents to climb internal leaderboards.

The pullback is the first structural signal that the per-prompt billing model Anthropic and OpenAI moved to in 2026 has broken the assumption that enterprise AI is cheap. Every token spent is now a line item a CFO can read. Every prompt an employee runs is now a metered event. The same shift that is forcing companies to control cost is also normalizing a layer of per-prompt behavioral tracking that did not exist when AI tools were sold on flat subscriptions. The companies cutting spend are also the companies that built the largest internal AI agent fleets, and the cost data they generate over the next two quarters will set the procurement default for every Fortune 500 buyer through 2027.

What the named companies did

Uber is the most concrete example. Andrew Macdonald, the ride-hailing company's president and chief operating officer, told a recent podcast that the company is now finding it "harder to justify" its outlay on AI tokens. "It's very hard to draw a line between one of those stats and 'OK now we're actually producing like 25 per cent more useful consumer features,'" he said. Uber has introduced usage caps that limit individual employees to $1,500 per month in token spending on individual AI tools, after the company blew through its entire 2026 AI budget by April.[1]

Walmart has taken the same approach on its internal AI agent. Suresh Kumar, Walmart's global chief technology officer, told the FT that use of the company's Code Puppy vibe-coding platform "really skyrocketed" and that Walmart is "now tak[ing] a step back." Employees are being tasked with identifying the right tools for tasks rather than reaching for the largest model by default.[1]

Cisco is balancing a desire to use the technology to deploy agents against the cost and availability of tokens. Jeetu Patel, Cisco's president and chief product officer, told the FT that "for every human you might have 10, 100 or on the aggressive side 1,000 agents . . . They just keep working and that consumes a chunk of [compute]." Patel added that "our engineers want more tokens . . . We have to figure out a way to fund it."[1]

Amazon warned employees last month that they should halt using AI "just for the sake of using AI" after engineers started to deploy agents for the sake of climbing internal leaderboards. The group has been forced to pivot its approach to measuring adoption in a bid to rein in costs attached to tool misuse. Meta took similar steps in April.[1]

Smaller companies are getting hit harder. Workato, a software company with 1,300 employees, saw its AI use explode after it put AI agents in the hands of staff last summer. "It took off like wildfire, people started really transforming their jobs with agents," said Carter Busse, Workato's chief information officer. Then Anthropic switched Workato over to token-based pricing in May. "Our spend went up 7x the first day and I'm like, oh shit, we created a monster," Busse told the FT. The "we created a monster" line became the FT's headline.[1]

The shift that broke the cheap-AI assumption

The structural cause is a billing model change. Throughout 2025, Anthropic and OpenAI sold enterprise access on flat subscriptions that priced AI access as if it were productivity software. In 2026, both labs moved substantial portions of their services to token-based billing, which tracks the units of data processed by models.[1]

The shift exposes companies directly to the cost of every prompt and automated workflow. Costi Perricos, Deloitte's global generative AI leader, framed the new model bluntly: "Compute costs are now beginning to enter the minds of both CFOs and boards. Consumers and businesses have been taught that AI is cheap or free and that is definitely not the case." Sam Altman, OpenAI's chief executive, said this month that cost had emerged as a "huge issue" for customers this year. "The issue never came up [last year] . . . People were totally happy with the amount they were spending."[1]

The shift is structural because it changes who inside a company can decide whether a query is worth running. Under flat subscriptions, an employee who pulled a model into a workflow had no marginal cost signal. Under token billing, every pull is a line item, and the people who see line items are the CFO and the board. The FT's central insight is that 2026 is the first year in which the enterprise AI market is being repriced in real time, with the new price visible to the people who write the checks.

The agent compute spiral

Agents are what made the bill spike. A chatbot takes a prompt, returns a response, and stops. An agent takes a prompt, decides what sub-tasks to run, runs them through a sequence of models and tool calls, and only returns when the task is complete. Each sub-task is a separate billable event. Cisco's Patel puts the per-human deployment ratio at "10, 100 or on the aggressive side 1,000" agents per human worker.[1]

Goldman Sachs analysts predicted last month that use of AI agents would result in a 24-fold increase in token consumption by 2030 and that the rise in demand would exacerbate a shortage of chips over the next 12 to 18 months.[1]

The structural read on the agent compute spiral is that the same companies preaching AI-driven productivity are now facing a bill that scales with how aggressively they push agents into workflows. The 24x token-consumption forecast is not a productivity gain forecast. It is a bill forecast. The companies whose productivity story depends on agent deployment are the companies whose CFO will see the bill first.

The privacy side: per-prompt billing is per-prompt tracking

The cost story has a privacy story underneath it that nobody at the named companies is talking about. Token-based billing requires per-prompt metering. Per-prompt metering requires that every prompt an employee runs be logged, attributed, and retained long enough to send an invoice. The shift from flat subscriptions to token billing is therefore also a shift from anonymous-aggregate usage telemetry to individually-attributed behavioral telemetry.

That telemetry is the same telemetry bossware vendors have been selling to enterprises for years. The 2026 bossware baseline is that 78 percent of companies now track their employees, with the surveillance extending to keystrokes, screenshots, location, and biometrics.[2] The new layer that 2026 token billing adds is finer-grained: not just whether the employee is at the keyboard, but what they asked the AI to do, what context they fed it, and what answer they got back.

The DataGrail 2026 report on shadow AI found that 63 percent of business software hides third-party AI subprocessors from customers, and 32.8 percent of AI systems handle high-risk data.[3] Per-prompt billing tightens that opacity: an employee using a sanctioned AI tool may think they are running a query locally, when the vendor's per-prompt meter is recording the query content, the model version, the response token count, and the user identifier at the same time. The OpenAI audited financials the FT published earlier this month show that the structural gap between AI revenue and AI compute cost is the gap that the per-prompt meter is closing in real time.[4]

The 2026 bossware story and the 2026 per-prompt billing story converge on the same data layer. Bossware vendors sell employers the right to see what employees are doing on the device. AI vendors sell employers the right to see what employees are doing with the model. The two converge when an employee runs an AI query that is metered by the AI vendor, attributed to the employee by the bossware vendor, and stored in a behavioral log that both vendors retain. The Workato 7x first-day spike is the visible cost of that convergence. The invisible cost is the behavioral log that the metering layer creates as a side effect.

China models overtake US on token consumption

One structural development the FT surfaces is a data-routing shift that maps directly to surveillance jurisdiction. Since the start of 2026, Chinese AI models have overtaken their US counterparts in token consumption on OpenRouter, an aggregation platform that lets users access multiple AI models.[1][5]

China's cheaper energy and more efficient models have allowed the country's AI labs to charge less than leading US groups for tokens, giving China a new edge on the AI battleground. The FT's chart of OpenRouter weekly token usage among the top nine models shows Chinese models overtook US models sometime in 2026 and the gap has widened since.[1]

The surveillance read on the routing shift is that the same enterprise AI buyers cutting their budgets in 2026 are also choosing where their prompts run. A token-routed to a Chinese model runs on infrastructure subject to Chinese data-access law. A token-routed to a US model runs on infrastructure subject to US Cloud Act and FISA 702 reach. The OpenRouter crossover means the per-prompt meter now spans both jurisdictions, and the price differential that drove the crossover is the price signal that enterprise buyers are following.

What comes next

Three watch points for the rest of 2026.

First, watch the per-quarter cost data from the named companies. Uber's April blow-out, Walmart's Code Puppy cap, Cisco's "figure out a way to fund it" admission, and Amazon's leaderboard retraction are all the first quarter of a multi-quarter experiment. The Q2 and Q3 earnings calls for these companies will be the first place the public sees the structural cost data. Investors who were sold on AI productivity gains will see what the productivity gains cost.

Second, watch for a second wave of companies cutting back. The FT story names Amazon, Walmart, Cisco, Uber, Meta, and Workato. Those are the early adopters. The second wave will be the mid-market enterprise buyers who signed contracts in the first quarter on the assumption that AI was a productivity line item and who now face a Q3 renewal at token-billing prices. Workato's 7x first-day spike is the canary.

Third, watch the per-prompt behavioral log as a regulatory surface. The EU AI Act will impose its first binding GPAI obligations on 2 August 2026, two weeks before the US fiscal Q3 starts.[6] Frontier AI labs serving the EU will have to publish training-data summaries and report serious incident statistics. Per-prompt metering data is the kind of telemetry regulators will eventually ask for in any incident investigation. The companies that built their internal AI cost-control on per-prompt metering are the companies that will be first to disclose what that meter is recording.

The FT story is a CFO story on the surface. It is a behavioral-telemetry story underneath. Both stories will play out in 2026.

References

  1. Financial Times via Hacker News and archive.ph mirror: We created a monster: companies rein in AI usage as costs strain budgets. Jamie John in London, Rafe Rosner-Uddin and Ryan McMorrow in San Francisco. 2026-06-19, reported on Hacker News as HN #48602571 (109 points, submitted by user fandorin at 19:57 UTC 19 June 2026). Archive.ph mirror: https://archive.ph/z24oE. Source data on OpenRouter cross-country token consumption: OpenRouter.
  2. State of Surveillance: 78% of Companies Now Track Their Employees. Here's What They're Collecting. 2026. Baseline corporate-surveillance context: 78% of employers monitor workers in 2026, tracking keystrokes, screenshots, location, and biometrics. The Microsoft Teams room-detection feature is the high-water mark for device-side bossware granularity.
  3. State of Surveillance: Shadow AI: 63% of Business Software Hides AI Models From You 2026. DataGrail 2026 report: 2,400 business apps tracked, 63% hide third-party AI subprocessors from customers, 32.8% of AI systems handle high-risk data.
  4. State of Surveillance: OpenAI Lost $38.5B in 2025. Your Data Pays the Gap. 2026-06-16. OpenAI 2025 revenue $13.07B vs total costs and expenses $34B, $20.92B operating loss, $38.53B net loss attributable to OpenAI; the structural gap is closed by capital raises, Pentagon contracts, and infrastructure subsidies. The per-prompt metering layer is the cost-to-revenue bridge the AI labs are now extending into enterprise procurement.
  5. State of Surveillance: DeepSeek and China AI Global Bans 2026. The 2026 wave of China-AI bans and procurement restrictions in the EU and US, the regulatory response to the OpenRouter token-consumption crossover.
  6. State of Surveillance: Norway Imposes Near Ban on AI in Elementary Schools 2026-06-20. The EU AI Act 2 August 2026 GPAI obligations deadline is the regulatory anchor for any Q3 disclosure of per-prompt metering data from frontier AI labs serving the EU market.