[DigitalToday reporter Choi Jae-won] Companies seeking to cut AI costs should do more than choose cheaper models. They need to reduce unnecessary tokens and computation, an argument said.
In an Oct. 1 article for IT media outlet TechRadar, Ben Canning (벤 캐닝), chief product officer at Alteryx, pointed out that as corporate AI use moves from experiments to large-scale operations, model usage costs are emerging as a new burden. Model routing and the use of open-weight models are being highlighted as alternatives, but he said those alone would be hard to solve the cost problem.
More recently, models have emerged that tout relatively low costs and high performance, such as Kimi K3 from China’s Moonshot AI, spreading a strategy of assigning appropriate models by task. But Canning argued that before deciding which model to use, companies should first determine whether a large language model is truly needed for the task.
If repetitive work such as file comparisons, applying internal rules and compliance checks is delegated to an LLM each time, the model reinterprets the same context and rules, repeatedly consuming tokens and computation. In particular, if the model has to newly grasp proprietary information each time, such as a company’s sales and margin calculation methods or regulations, not only can costs rise but the risk of wrong answers can also increase.
The alternative presented was “business logic layers.” Companies would build analysis procedures for repetitive tasks and company-specific rules and definitions into workflows in advance, and then have the LLM call up and explain the results rather than do the calculations itself. For example, if a team’s margin is asked, completing the calculation on a cloud data platform and having the model deliver only the result can reduce large-scale context processing and duplicate computation.
In fields such as tax, finance and compliance, where consistent answers are required under set rules, such a structure is important for accuracy as well. Canning stressed that the key to optimising AI costs is not finding cheaper tokens, but building reliable workflows and business logic so that unnecessary tokens are not generated in the first place.