The Growing Cost of Intelligence (Part I)
Why costs are rising, why they matter, and how enterprises should respond
This is Part I of a two-part series on rising AI costs. Part I covers why costs are rising, why they matter, and how enterprises should respond. Part II will explore what this means for investors.
AI was supposed to make intelligence cheaper.
But instead, it’s becoming one of the fastest growing line items on the P&L.
Take Uber for example. The company blew through their annual token budget in a single quarter (and is now capping token spend at $1,500/month). Another company reportedly blew $500m on Claude AI in one month due to no usage limits on licenses for their employees.
In the valley, there’s a new phrase for this type of behavior: “tokenmaxxing”.
It refers to the practice of maximizing AI token spend, often treating token consumption itself as a proxy for productivity. In some companies, AI usage is even being scored on leaderboards, a way to prove who is most “AI-native”.
Tokenmaxxing may sound like a meme, but it is just a reflection of what we already know about history. When a new technology suddenly expands what is possible, the first instinct is rarely optimization. It’s experimentation. Teams try everything. Adoption becomes indiscriminate and unfettered. AI is added to every possible workflow and absolute usage becomes a sign of progress. Only later, does optimization matter.
We’ve seen this before - the internet and mobile made distribution cheap, and companies later had to relearn retention and CAC discipline. Cloud made compute easy to provision, and companies later needed FinOps to control spend.
AI is entering its own optimization phase now.
Our hypothesis is that the next few years of AI will be defined increasingly by optimization and efficiency.
The reason is simple: AI is cheap in proof of concept, expensive in adoption, and even more expensive when it becomes infrastructure. And as AI moves from experiment to infrastructure inside the enterprise, costs become harder to ignore.
That shift will not only shape how enterprises adopt AI, but will shape what types of AI startups get built and funded, and where AI researchers focus their energy.
Why are costs so high?
At first glance, this seems counterintuitive.
Unit token costs have come down dramatically. Model providers are improving efficiency. Open-source models are getting better. Competition is increasing. Yet, AI costs are exploding.
AI companies are already adapting to this reality. Anthropic is a good example. The company recently shifted Claude Enterprise toward usage-based billing on top of a monthly seat fee. Salesforce and ServiceNow have moved toward consumption-based pricing for their AI products as well.
This is not just a new pricing strategy. It’s a reflection of a market where demand is far outstripping supply.
The demand side: workloads are changing
AI workflows are no longer linear. They are iterative by design.
This is easiest to see in coding agents. Tools like Claude Code and Codex are not one shot systems. A developer does not make a single request and move on. The workflow becomes a loop: it generates, observes, evaluates, learns, and repeats.
What used to be one operation in traditional software becomes many operations in an agentic workflow.
This is the hidden multiplier behind AI demand.
As intelligence improves, we ask more of it. We add more context. We introduce more steps. The system becomes more capable, but also much more compute-intensive. Cheaper tokens doesn’t matter if each workflow consumes 100x or 1000x more of them.
The same dynamic is showing up in training and post-training. Inference is no longer just part of serving a model to users. Increasingly, it becomes part of the training process itself. Models generate outputs, evaluate outcomes, and recursively improve through repeated attempts. In other words, the path to better intelligence requires more intelligence to be generated in the first place. This self-improving loop is what makes AI powerful. But also what makes it expensive.
The supply side: the physical world is the bottleneck
On the supply side, the bottleneck has moved from the chip level to the systems level.
Increasingly, the constraint is not just the GPU. It’s also everything around the chip: memory, networking, cooling, data center equipment, electrical transformers, and the physical materials required to turn chips into tokens.
The problem is no longer one bottleneck. It is a chain of bottlenecks, and each delay compounds the next.
Labor is becoming a bottleneck as well. There simply are not enough electricians and skilled tradespeople to build and maintain the physical infrastructure AI now requires.
And then there is power itself.
Data centers are becoming one of the largest new sources of electricity demand. Yet U.S. power generation has been largely flat over the past 15 years, and shifting energy policies across administrations have only added to the uncertainty. Even when new power can be brought online, the grid was not built to absorb this much load this quickly.
Local pushback is also increasing. Across the U.S., fourteen states are considering bans or moratoriums on new data centers. The result of all this is that half of U.S. data centers planned for this year are expected to be delayed.
This is the core tension behind AI today:
Demand scales like software but supply scales like physical infrastructure.
A model can improve in days. A new AI workflow can be designed overnight. An agent can spawn thousands of reasoning steps in seconds.
But the supply side does not move in that way. Electrical components require months of lead time to procure. Data centers take years to finance, build, power, and connect. Legislation and permitting move even more slowly. None of this scales at the speed of software.
This mismatch is why AI costs will remain stubbornly high in the short-to-medium term. And it’s why cost optimization will be increasingly important in the next few years. When infrastructure cannot move at the speed of software, optimization becomes the only way to keep scaling.
Why This Matters for Enterprises
For Enterprises: When tokens substitute labor
For the last three years, AI spend was mostly treated as an innovation budget.
A few POCs here or there. The spend was easy to justify because it lived in the experimentation bucket.
That’s starting to change.
AI spend is moving out of one-time innovation budgets and into regular cost centers. It’s becoming part of the permanent operating model. And once something becomes permanent, the ROI must become clearer.
But there is a second reason why cost is starting to matter more: AI is beginning to substitute for labor. But unlike labor, AI is far easier to measure and benchmark.
Let me explain.
In the old world, labor costs were messy and subjective. If one company spent more on headcount than another, there were many ways to explain the difference: better talent, more experience, different geographies and cost of living, etc. People also bring harder-to-measure qualities like creativity, communication, and team chemistry. There’s a lot of subjectivity in how much a person’s worth is and how much they should get paid.
In other words, corporate inefficiency had plenty of places to hide.
But AI creates fewer hiding spots.
If two companies use similar agents to automate similar workflows, but one spends twice as many tokens to get the same result, that gap becomes much harder to explain away. Cost per ticket, cost per line of code, or cost per contract reviewed, becomes visible in a way labor productivity often was not. A unit cost of intelligence becomes much easier to compare.
This is where the next phase of AI adoption gets more disciplined. Boards will increasingly ask management teams to prove the ROI of AI spending and benchmark their spend against peers.
The next phase of enterprise AI adoption will be more disciplined because the cost of intelligence will be much easier to mesaure.
The Enterprise Lens
How Should Enterprises Organize Around AI?
The first question to managing cost is deciding who owns AI.
Is it the CIO, CTO, CFO, CDO, or a newly created Chief AI Officer? Should AI decision-making be centralized, or distributed across individual business units?
At first glance, there is a strong case for centralization. A centralized AI function can standardize model access, negotiate vendor contracts, enforce security policies, monitor usage, and prevent every team from rebuilding the same infrastructure. It reduces duplication, improves governance, creates purchasing leverage, and gives leadership a clearer view of where AI dollars are going.
Over time, centralization should reduce the overall cost.
But centralization has a hidden cost too: distance from the actual workflow.
Centralized teams rarely understand every business process deeply enough to redesign it from the inside. The highest-value AI use cases often live closest to the work itself: engineering, customer support, sales operations, legal, etc. If every use case has to wait behind a central team, experimentation slows. The company risks optimizing for cost before it has even discovered where AI creates value.
This is where AI adoption looks very different from traditional SaaS adoption.
With SaaS, adoption was often top-down. If the CEO or CFO decided the company was going to use Salesforce as the system of record for sales, everyone was expected to learn and adopt it. The workflow adapted to the software.
In AI, it is the opposite. The software is expected to adapt to the workflow.
That makes adoption inherently more bottoms-up. Marketing teams choose their own agents. Engineers decide between Codex, Claude Code, or Cursor. Legal teams can experiment with different legal copilots. Individual users and teams often feel the productivity gains first, so they have the strongest incentive to experiment.
That argues for more decentralization than in the SaaS era.
But decentralization has real costs too. Every team buys its own tools, runs its own pilots, selects its own vendors, and builds its own agents. Usage sprawls. Spend duplicates. Security surface area expands. Leadership loses visibility into cost, governance, and ROI.
So the answer is not full centralization or full decentralization. It is governed decentralization.
Centralize the things that benefit from scale: model access, procurement, security, governance, observability, cost tracking, and shared infrastructure.
Decentralize the things that require workflow context: use case selection, process design, and day-to-day ownership inside the business units.
In other words, the center should build the rails. The business units should decide where the trains go.
Tactically, this could mean standing up an AI Center of Excellence (CoE) chaired by a senior executive, but staffed with representatives from each major function. The CoE’s job should not be to approve every use case. It should manage the shared layer while giving business units enough autonomy to move quickly within the workflows they understand best.
That is how enterprises get the cost discipline of centralization without losing the speed and context of bottoms-up adoption.
Buy or Build?
The buy-versus-build question in AI should come down to three questions:
How AI-forward are we as an organization?
How important is the workflow to the business?
Who can improve the cost curve faster: us or the vendor?
Start with the question: how AI-forward are we?
For companies that are not AI-native, building internally can be more expensive than it looks. The spreadsheet may make building look cheaper, but the cost is speed.
AI is moving so quickly today that a six-month internal build can become stale before it launches. By the time the team ships, the model landscape may have changed, agent frameworks may have matured, and vendor products may have advanced one or two generations. In a market moving this quickly, building is not always cheaper. Sometimes it is just a slower way to recreate yesterday’s product.
So for companies that are not AI-forward, the default should usually be to buy first – but buy in a way that builds internal capability over time.
That means choosing vendors not only for their product, but for how much they help the organization learn. Prioritize vendors that are more flexible and can work closely with the internal teams to co-design workflows, train internal teams, and build the internal team’s AI muscle.
This is one reason startups can be attractive partners. Large vendors may offer broader platforms, but smaller companies are often more willing to spend time with the customer, customize around specific workflows, and transfer know-how. For enterprises trying to become more AI-capable, this learning may be as valuable as the product itself.
The second question to ask is whether the workflow is core to the business.
If the workflow is not central to differentiation, buying is often the right answer. There is no need to build a non-core functionality if a vendor can solve it faster, better, and at a reasonable cost.
But if the workflow is core, the answer changes.
In his book Delivering Happiness, the late Tony Hsieh wrote about Zappos’ decision not to outsource customer support, even though outsourcing would have been far cheaper. His reasoning was simple: customer support was not a back-office function. It was the brand itself. The call center was not a cost center to be minimized - it was core to the company’s identity.
AI should be evaluated the same way. If something is central to your differentiation, you may not want it to be managed by a third-party vendor. Buying may still be the right starting point – it gets you speed, learning, and a baseline product. But over time, the company should quickly build toward more ownership and control.
The third question is who can improve the cost curve faster: you or the vendor.
Buying gives speed, but it also ties the company to the vendor’s architecture. The customer inherits the vendor’s model choices, product roadmap, infrastructure decisions, and margin structure.
That may be fine early on. But over time, the economics depend on who can improve the system faster.
If the workflow generates meaningful learning from your own users, customers, data, and internal processes, you may be in the better position. Each piece of feedback can reduce context, redesign the workflow itself, and lower the cost.
The vendor may have broader usage across many customers which has learnings as well. But the company will have the richer context-specific learning.
So the real question is: who can turn feedback into a better cost curve faster? If the vendor has the learning advantage through its scale, buying may remain the right answer. But if your own workflow creates the feedback needed to improve quality and reduce cost, ownership becomes more important.
Taken together, these three questions point to a conclusion: buy-versus-build is not a static decision. It should be treated as a staged path.
Buy where speed matters. Build where control matters. And over time, continually reassess the tradeoff as your AI capabilities improve, the workflow becomes more strategic, and your own feedback loops begin to change the cost curve.
Tactical recommendations for enterprises to control costs
The next phase of enterprise AI will not be defined only by who adopts AI fastest. It will be defined by who manages it best.
Here are some tactical recommendations for enterprises to consider:
Create visibility before enforcing control.
Most companies still do not have a clear view of where AI spend is going. Usage is scattered across model providers, copilots, SaaS vendors, API calls, internal agents, and individual employee subscriptions. Before companies can optimize anything, they need to know which teams are using AI, what workflows they are applying it to, and how much each workflow costs.
But internal visibility is only half the picture. As AI begins to substitute for labor, companies will also need external visibility: how their AI cost per workflow compares with peers and competitors. If two companies use similar agents to complete similar work, but one spends twice as much to get the same outcome, that gap needs to be addressed right away.
The right metric is not cost per token. It is cost per unit of work: cost per ticket resolved, cost per line of code pushed into production, cost per contract reviewed, cost per customer interaction, or cost per workflow completed.
Separate experimentation process from production process
AI experimentation should remain bottoms-up. Each team should have a small dedicated budget to test tools, try new workflows, and discover where AI creates value. The people closest to the work are usually best positioned to identify where AI can help.
But once usage moves beyond experimentation, it should go through a more centralized process. At that point, the company needs to evaluate vendor overlap, security, pricing, usage limits, integration requirements, and whether the workflow is important enough to become part of the permanent operating model.
This distinction matters because many AI projects look cheap in POC but become expensive at full-scale production. A small pilot can quickly turn into a recurring cost line once more users scale across the company.
Ultimately, the goal is to preserve bottom-up discovery while centralizing scale. Teams should have the freedom to experiment within clear budget limits, but broader deployment should require a higher bar: clear ROI, cost controls, and a decision on whether the company should buy, build, consolidate, or shut it down altogether.
Redesign the workflow, not just the model call.
The biggest savings will not come from making every prompt slightly cheaper. They will come from changing the workflow itself.
Many companies will start by adding AI on top of existing processes: summarize this meeting, draft this email, answer this support ticket, review this contract, etc.. That helps, but it often preserves the same underlying complexity. AI makes the existing workflow faster, but not necessarily better.
The bigger cost opportunity is to redesign the workflow from scratch. If AI can resolve a support issue before it becomes a ticket, the cost savings are greater than simply making ticket responses cheaper.
This is similar to Elon Musk’s operating philosophy at Tesla. His process is to first question every requirement, then delete the unnecessary part or process, and only then automate. The sequencing matters. Automation should be the last step, not the first.
That lesson applies directly to AI. Many companies will be tempted to use AI to automate every existing process. But if the process itself is unnecessary, fragmented, or poorly designed, AI simply makes the waste run faster. One of the biggest mistakes is automating a process that should have been deleted in the first place.
Assign a business and technical ownership to each AI workflow
Every production AI workflow should have both a technical owner and a business owner.
This matters because AI workflows sit between software and operations. They are not just systems to be maintained, and they are not just business processes to be redesigned. They are both. If ownership sits only with the technical team, the company may optimize the model without changing the actual workflow. If ownership sits only with the business team, the workflow may gain adoption without the right controls around reliability, evaluation, security, data access, and cost.
This shared ownership is especially important because AI is moving so fast. Models change, prompts drift, data sources evolve, users find workarounds, and edge cases appear over time. A workflow that looks successful in a pilot can degrade in production if no one is continually monitoring it from a business outcome and systems perspective.
In practice, every production AI workflow should have a clear answer to three questions:(1) who owns the system (2) who owns the business outcome and (3) what metric(s) defines success.
Route work to the cheapest system to do the job
Not every AI task requires frontier intelligence. In fact, not every task requires an LLM at all.
The key is to separate reasoning from execution. Use transformer-based technologies (LLM, VLMs, etc.) where judgment, ambiguity, language understanding, or planning is required. But once the system knows what to do, execution should be as deterministic as possible: APIs, rules, and traditional software.
This is especially important as enterprises adopt more agentic workflows. For example, a vision-based computer-use agent may be useful when no structured interface exists, but it is often a much more expensive way to automate a workflow that could be completed through direct APIs.
Some workflows will still need the best reasoning model available. A customer-facing workflow, complex coding agent, or regulated decision process may justify frontier intelligence. But many tasks can be handled by smaller models, open-source models, deterministic systems, or no model at all.
That requires more than an understanding of model routing, prompt optimization, caching, and evaluation frameworks. It requires a deeper understanding of the workflow itself: which steps require reasoning, which steps require software execution, which steps require human approval, and which steps should be removed entirely.
Turn proprietary data and workflow into a learning flywheel.
The best AI systems should get cheaper and better with use. But that only happens if enterprises capture the data from real workflows: what people ask, what context was needed, what actions were taken, where the system failed, and which final outcomes were accepted.
Over time, this creates a flywheel. More usage creates more workflow data. Better workflow data improves retrieval, model routing, and workflow design. That leads to fewer retries, better context management, less human intervention, and lower cost per outcome.
That is how AI cost advantage becomes defensible. Two companies may use the same foundation models, but they will not have the same proprietary workflow data - or the same ability to turn that data into cheaper, smarter systems.
Conclusion
AI was supposed to make intelligence cheaper. In the long run, it probably will. But in the near term, the opposite is happening: as AI becomes more capable, companies are finding more ways to use it, and every new workflow creates more costs.
That does not mean enterprises should pull back. It means they need to become more deliberate. The next phase of AI adoption will not be defined by who uses the most tokens or launches the most agents. It will be defined by who can turn AI into measurable productivity - and who can make the systems cheaper over time.
In Part II, we will explore what this means for investors: what does the “cost optimization stack” look like, where are the new opportunities, and which startups are best positioned to benefit.
Stay tuned.









