📣 Send us your press release
Site updates every 15 minutes
Technology

AI Model Benchmarks Don't Predict Actual Costs

Alibaba's Qwen 3.8-Max and Anthropic's Claude Opus 5 highlight how raw benchmark scores can be misleading. Time and token budgets significantly impact performance and actual costs.

6 August 2026
AI Model Benchmarks Don't Predict Actual Costs

Raw performance metrics for artificial intelligence models do not always accurately reflect their real-world operational costs. Recent tests involving Alibaba's Qwen 3.8-Max and Anthropic's Claude Opus 5 have revealed that model performance can vary dramatically based on the time and token budgets allocated to them.

Alibaba initially positioned its Qwen 3.8-Max model as nearly on par with Claude Opus 5. However, independent tests using different time and token budgets yielded conflicting results. When Qwen 3.8-Max was given considerably more time and a larger token budget, its performance improved significantly. This illustrates how a model's "best effort" can differ substantially from its results under default settings.

Experts advocate for a shift from solely relying on performance metrics to a cost-per-successful-task model. This approach calculates the total cost, including failed attempts, and divides it by the number of tasks successfully completed. Concurrently, time and token budgets should be clearly defined as part of the acceptance criteria.

Traditional pricing, such as price per token, is no longer sufficient for predicting actual costs. Models that expend a significant portion of their token budget on reasoning may reach their token limit before producing an output, leading to an empty result at the cost of a full run. This is particularly problematic when comparisons between models are based solely on published prices.

In conclusion, when evaluating the true costs of AI models, it is crucial to consider the impact of time and token budgets. Cost per successful task and clear budget criteria offer a more accurate representation of a model's practical value than performance benchmarks alone.

Original source: venturebeat.com