Search results for Speculative Decoding
AI & Enterprise
OpenAI says GPT-5.6 Sol optimises GPU efficiency itself, cuts inference costs 20 percent
OpenAI says it used its latest AI model, GPT-5.6, to optimise its own AI infrastructure, lowering inference costs and improving throughput. It said improvements to the inference engine and GPU use cut end-to-end service costs by 20 percent for the GPT-5.6 Sol model and raised token generation efficiency by 15 percent. OpenAI also optimised KV cache handling, GPU workload distribution and agent operations, limiting tool output and improving caching.
AI & Enterprise
Cerebras compresses 163 seconds into 5 seconds, says GPU era is over
AI chip designer Cerebras put the 1 trillion-parameter open-weight model Kimi K2.6 into its enterprise inference service and achieved 981 tokens per second, a pace it says is the world’s fastest. It also cut the time to complete 500 output tokens from a 10,000-token input to 5.6 seconds, versus 163.7 seconds on the official Kimi endpoint. The company is pursuing an IPO and reported 2025 revenue of $510 million and net profit of $238 million.
AI & Enterprise
Red Hat unveils Red Hat AI 3.4 with up to three times faster inference, expands space and auto ties
Red Hat announced new products and partnerships across enterprise AI platforms, space computing and software-defined vehicles at its annual Red Hat Summit 2026 conference on May 11. It highlighted the Red Hat AI 3.4 update, adding model service capabilities and speculative decoding to boost inference speed by up to three times. Red Hat also expanded cooperation with Nvidia, announced a deployment on an ISS edge micro data centre with Voyager Technologies, and unveiled a joint initiative with Nissan.