Nobody Can Tell You What An AI Task Costs: Not Even The AI
Konstantin Klyagin is the Founder of Redwerk and QAwerk, driving innovation in custom software development and quality assurance since 2005.
gettyOver the last year, we’ve seen top companies shift from leaderboards that track AI use to capping AI tokens per employee. Both extremes miss the mark. We now have solid proof that the true cost of an AI task is impossible to predict reliably. So, how do you protect yourself from massive, unexpected AI bills?
The answer isn’t just tracking the dollars you spend on tokens. In this article, I’ll give you an honest breakdown of how to calculate the real value of your AI use, so you can make the best decisions for your business.
First, let’s recap the tokenmaxxing hype that has just died down. We have some very good examples here:
Amazon ran an internal leaderboard to rank employees on token consumption. This led people to use AI for every trivial thing, and computing costs shot right into outer space. Uber burned through the 2026 AI budget within four months. The company’s course correction was to cap tokens at $1,500 a month per employee (per tool), adjustable for some engineers. Microsoft canceled most internal licenses for external coding agents and moved thousands of engineers to its own tooling.
The message is clear. Companies went from using AI for the sake of using AI to frantically budgeting its use. However, this solution might be just another trap.
Here’s why. Measuring your AI value in dollars is just as faulty as measuring it in consumed tokens. Let’s look at OpenAI to see this in practice. It recently published a report on research acceleration within its labs. According to the data, they see an eight-hour workday of a researcher’s labor as equal to about 3.1 agent-workdays of research effort. However, the same report also states that over half of successful tasks that run four to eight hours require human intervention.
That single line takes us from the clear-cut math of an AI agent doing 3.1 times as much as a human researcher in one day to the uncertainty of completion.
How can you compare these things if you aren’t sure how many human hours will actually go into verifying and correcting what the machine has done? How many interventions will be needed for each specific task? And finally, how many tokens will be spent on that task and all the necessary corrections?
Now we get to the root of the issue and why I believe setting a dollar cap on AI use isn’t efficient if you’re chasing productivity. You can easily run this kind of experiment yourself. Give your AI agent the same task twice and look at the number of consumed tokens. That number might differ a bit or a lot, but it will always differ.
• Agentic AI tasks are much more expensive than code reasoning and code chat (up to 1000x).
• Token usage can vary by up to 30x on the same task.
The second point is the core reason why AI budgeting doesn’t really do much. That same study also showed that higher token usage doesn’t translate into higher accuracy. This means AI models sometimes use more tokens with no reasonable explanation for you as the end user. And your CFO must budget for that.
I can see this variance in the cost per delivered feature directly in my own business. It’s quite different from what we’re used to seeing with the set cost per seat in traditional software. However, the switch to agentic AI is our reality now, and SaaS is already becoming a thing of the past. So our goal is to learn how to maximize AI budgets, not ration them and ruin productivity gains.
Cheaper models won’t help. They certainly affect budgets to some degree. However, token cost means little when you’re measuring the wrong metric.
What I believe is the main change we need to make as business owners in our approach to AI is to focus on its actual value. In business terms, this means we must start measuring outcomes, not token use or dollar spend. Here’s how I’d go about this:
1. Define the completed outcome first. You need to know exactly how much it costs you to complete a specific task, for example, merge a code change or resolve a ticket. You need data on how much time and money it takes with and without AI. Measured from the moment you get the outcome.
2. Make the cost visible. I like how Uber lets engineers see the running cost on their terminals. This makes the actual spend real for users and removes the element of surprise of getting the AI usage bill at the end of the month.
3. Cap the mechanism. We implement methods such as context window limits, longer cache windows and rerouting specific tasks to cheaper models. This way, AI budgeting moves from a FinOps-exclusive problem to an engineering one.
4. Track the spread, not the average. Remember that 30x token use variance? Average numbers can’t be trusted in this sector.
Looking at how top global businesses are fighting the tokenmaxxing problem today, we can see an emerging pattern. The real solution isn’t to limit the budget but to understand the value AI delivers within your processes.
Some businesses can minimize spend, but others will still have huge AI bills. The difference is whether that bill makes sense in your unique situation, and it’s something only you can calculate by measuring real outcomes.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

