In Q4 2025, Claude Code got so good that coding agents took off among normies. Even technically unsophisticated users like myself could extract huge leverage out of them. That led to 2 things:
Coding agents as daily active user (DAU) products for non engineers
Token centric/AI native/inference heavy workloads getting into production scale for non-intelligence companies.
I observed both of these personally as a) thats when I became a DAU and b) thats when Substrate’s Claim Status agent first deployed in a stable way in enterprise. I think this shift kicked off the #tokenmaxxing (or as someone recently put it to me “token guzzling”) phase.
That era is now firmly over.
If you’re a CFO at any real scale and responsible for protecting and extending a company’s free cash flow, you’re looking at model spend, and requiring at least as much productivity/revenue lift to justify it. If you’re not selling intelligence, your revenue didn’t also go vertical in Q2. Given this, not surprising the tokenmaxxing era lasted exactly for a single quarter.
The consensus view over the last few years was that neoclouds (Eg CoreWeave) and inference clouds (eg Fireworks/BaseTen) would be fine businesses, but would play second fiddle to the frontier labs, because frontier, closed source models would always be the premium choice. Today this view looks far more endangered, for 2 primary reasons
Frontier model cost has ballooned enough that most companies are enforcing some form of cost control (either capping token spend, bounding usage to subscriptions, or rotating workloads to open source)
Open source and open weight models have dramatically improved (and everything that’s caught up to Opus 4.7 is good enough to do most run of the mill enterprise workloads)
The tokenmaxxing era, short lived as it is, helped expose a bunch of problems that companies are grappling with now:
Spending Blind: Most businesses that saw inference spend go vertical in Q2 2026, mostly have no idea what it was spent on. Even then, almost no one expects token consumption to drop.
Token optimization and routing is here: The most technical companies are already heavily optimizing token spend either through in house means or through routers (I wrote about routing layers here a half decade ago - I continue to believe that for any mission critical systems, if you’re at real scale, you need to have a optionality in your supply chain (ie a router) in place. Intelligence is no different: https://www.kunle.app/may-2020-API-routing-layers.html). But if your primary product is not intelligence, or you’re not at enormous scale, building custom routing infra for yourself isnt rational.
Frontier models get deprecated annually or 2x/year: Intelligence companies have had to deal with this problem for a couple of years already, but their teams are heavily staffed with engineers and researchers who understand these problems in a nuanced way. If you’re not one of these companies, the odds are you’re not staffed to have someone step off production to rotate to a new model every time a frontier lab deprecates one under you.
Model selection is not a one time decision: The first burst of non-intelligence companies becoming dependent on inference probably happened during H1 2026. Those companies now have use cases that are already stable, and high performing, on a model that probably wont exist in a year. When those models get deprecated, these companies will be forced to migrate, which means they all have to have roadmap budget and capacity simply to play defense. Think about what this means;
If you built a product on a frontier model that works well
You’ve already optimized across cost, latency, and experience.
You get this email
Now some of the most qualified members of your team has to migrate, and spend time just to maintain a configuration that was working.
Companies need (at least) two different things from models
For some tasks, they need frontier intelligence, so the best model should keep changing over time. Cyber security is one such category of task - as the frontier moves, you take on catastrophic risk by not moving with it. Others might include fundamental research in your domain with hit dynamics eg blockbuster drugs, breakthrough materials science etc. For other tasks, enterprises need determinism and stability, so once the product works, they should anchor to a controlled model setup and optimize cost and experience.
Downstream, any task that needs determinism and stability (and needs to be low cost) ultimately needs to be rolled to an open source or open weight model, and ultimately on a cloud you own or hardware you control. If you are on a frontier model at some point, it will be deprecated. Your team will be back spending time trying to give clients an experience that they’ve already had, that you already were reliably servicing, and that is now going away.
Imagine if, once a year, Stripe changed the payment infrastructure that you were relying on, and random new credit cards would start failing.
Large-scale software and AI companies already have this because it’s the primary purpose of theirs to sell intelligence, or intelligence is existential for them. Coinbase is an example of this:
You would assume that a company like Abridge or Open Evidence or Harvey or Legora (companies that sell intelligence but are not frontier labs) would have people doing this natively all the time because it’s their job, and that’s true. Coinbase is a public company, and it doesn’t sell anything that’s dependent on inference, yet the spend required them to do this.
Routing and evals are the same problem
We think of routing today and evals as separate problems, but I believe that they’re really the same problem expressed in different ways.
Model routing tools today optimize primarily for cost. They’re helping you pick the cheapest tool for the job.
Eval’s tools today optimize primarily for experience or reproducibility.
Those tools only sometimes explicitly handle non cost/quality dimensions:
latency,
throughput,
variance,
failure mode severity,
auditability/compliance,
downstream business outcomes
margins (assuming your revenue is also denominated in inference)
Both of these are solving for “which model I should use right now for this task”. Because neither routinely optimize the remaining dimensions, you frequently require a human to optimize those.
Even then, this optimization is at the level of a workload (ie a whole set of tasks that are encountered and run repeatedly). Dynamic optimization at the level of a workload or even better, a task, in a way that is aware of these other dimensions, still falls to a human, and lots of companies are inadvertently building this in house.
This layer is simultaneously aware of evals (a hard proxy of your product outcome, or stated differently, your internal benchmark), customer feedback (your customers’ actual experience), product pricing and costs (to assure margins) and can influence all of them.
In some ways, you can think of this as a private benchmarking tool. For research companies, benchmarks are relevant because it’s one of the ways that they get graded on by the market at large. For non-research labs, the concept of benchmarks is deeply challenged, because very often benchmark performance doesnt translate directly to business outcomes for production workloads (or at least it takes a gigaton of work for that to be true). What enterprises need specifically is: how well does this model perform against my specific tasks in my specific environment? What changes in harness/fine tuning/post training/RL does it take to get it there?
This tool would watch success rates, test alternative models, automatically test new models against your own evals when released, and proactively
identify the cheapest model that preserves required performance
Alternative models and harnesses and their various tradeoffs for speed/latency and experience,
enable load balancing (eg for outages) and to take advantage of burst pricing
If you knew that multiple inference clouds and frontier labs had a model that just about amounted to the same cost-quality trade-off for you, you could almost batch your traffic on a continuous basis and await burst pricing deals to really optimize economics on your token spend.
As more companies have costs and revenue denominated in or influenced by tokens and inference spend, this becomes more and more required. Beyond that, the deeper opportunity is a company that helps you optimize product and business performance with models as an input. This can start as routing based on cost at the workload level, but ultimately what you want is really, really fine-grained control and routing at the task level.
For what it’s worth, my gut is the research labs and the large-scale inference-heavy companies have always had this problem and actively prosecuted it, but token spend and model deprecation have brought this to the fore for everyone else.
Things I don’t understand yet
It’s likely that Deepseek or Moonshot (or some other open weight lab) will launch a frontier model that meaningfully beats the closed source/US frontier labs on a material axis beyond cost. Stated differently, it is likely that in the next year, the leading model available to the public will be an open weight one, at least for a time. When that happens, it will be very expensive to not have your routing figured out. At that moment, those who have their routing figured out will much more quickly be able to deploy the best experience in the world (for their use case) than those who do not.
The last thing that I havent accounted for is some combination of models, coding agents, and IDEs might become so good that transitions and migrations are effectively “free”. This feels crazy to me given the prevalence of subtle and nuanced errors that emerge whenever you make large shifts in complex systems, but given the rapid advancements in coding agents over the course of this cycle, you can’t rule it out. If this happens, the net cost of maintaining this infrastructure within a company will dramatically drop, and so the company-size, scale or size of the engineering team you will need to have in order to do this in-house might shift as well.





I love this write-up