AI & Enterprise
OpenAI says GPT-5.6 Sol optimises GPU efficiency itself, cuts inference costs 20 percent
OpenAI says it used its latest AI model, GPT-5.6, to optimise its own AI infrastructure, lowering inference costs and improving throughput. It said improvements to the inference engine and GPU use cut end-to-end service costs by 20 percent for the GPT-5.6 Sol model and raised token generation efficiency by 15 percent. OpenAI also optimised KV cache handling, GPU workload distribution and agent operations, limiting tool output and improving caching.