What Happened
AI spending is becoming irritatingly difficult for businesses to forecast, the WSJ argues, because the task completion cost can vary widely across models and even across repeated runs of the same model.
So far, so mostly unarguable. The WSJ provides. as further support, a survey in which only 11% of companies forecast AI spending within 10% of actual costs, and showed that lower-priced models sometimes cost more overall than higher-priced ones.
The core of the piece compares two Google models performing the same task. Gemini 3.1 Pro, a frontier-ish model, completed it successfully in 85 steps for roughly $1, while the cheaper Gemini 3 Flash took 952 steps, failed, and it cost $14 in tokens.
The article concludes that businesses cannot estimate task cost from model pricing alone. Model selection, task complexity, retries, failure rates, and token consumption all matter.
What It Means
This is all fair enough, but it is much less striking than the WSJ thinks it is. And, whether intentionally or not, it carries water for frontier AI companies: their salespeople will have been spamming customers non-stop for the last 24 hours: CHECK THIS WSJ ARTICLE OUT email subject lines are filling inboxes worldwide.
Sure, AI spending is much more variable than traditional software spending. Agentic systems can retry, wander, invoke tools, consume unpredictable amounts of context, and fail after doing substantial work. That makes individual task costs noisy and enterprise budgeting harder.
But the piece goes further by giving too much weight to unusual cases in which cheaper models become more expensive because they take far more steps. That is not the normal pattern as models get better. Outside of genuinely weak or badly matched models, lower-cost models remain lower-cost for almost all tasks for which they are suited.
