A study found low-priced AI models are not always efficient. [Photo: Reve AI]

A study has found low-priced artificial intelligence (AI) models are not always cost-efficient.

Gigazine reported on Oct. 6 that research teams from Stanford University, Carnegie Mellon University, the University of California, Berkeley and Microsoft (MS) Research had eight state-of-the-art AI models perform more than 6,800 tasks. They said the cheaper model’s total cost was higher than the more expensive model in 32 percent of cases.

The study compared how lower-priced and higher-priced models actually structure costs across tasks in areas such as math, programming and science. The researchers said it is hard to explain real spending using only surface-level token prices or per-call costs. That is because the number of reasoning attempts and action steps needed to finish the same assignment can vary widely.

As an example, the researchers cited a task that analyses a YouTube video frame by frame. Gemini 3.1 Pro, a higher-priced model, finished in 85 steps. Gemini 3 Flash, a lower-priced model, failed to finish even after more than 1,000 steps. The task cost $1 for Gemini 3.1 Pro and $14 for Gemini 3 Flash.

The researchers pointed to AI “overthinking” and “overacting” as reasons for the cost reversal. They said complex tasks can easily require more reasoning and execution steps, and higher-performing models can sometimes finish faster and at lower cost. In tests across broad fields, they also found a case in which Gemini 3 Flash consumed more than 60,000 thought tokens, while higher-performing GPT-5.4 solved the problem with 25 tokens.

The study also showed it is hard to predict costs. When the researchers repeatedly entered the same instruction into the same model, there were cases in which the most expensive attempt cost 9.7 times more than the cheapest attempt. That suggests large spending volatility not only in model selection but also in repeated runs of the same model.

The findings have prompted calls for companies to move beyond simple unit-price comparisons in their AI adoption strategies. Lower-priced models can be advantageous for simple work, but total execution costs can rise for complex tasks. Linzhao Chen, who conducted the experiment, said, “You should not judge which model is cheaper just by looking at price.”

The study reaffirmed that price lists for AI services do not directly reflect actual operating costs. For tasks that require multi-step reasoning and tool use, performance gaps between models can directly affect total cost. Companies that adopt paid AI services may need to review task completion rates, the number of steps and variations in repeated runs, as well as per-input pricing.

Keyword

#Stanford University #Carnegie Mellon University #University of California Berkeley #Microsoft Research #Gigazine
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.