Search results for Speculative Decoding
AI & Enterprise
OpenAI passes 1 billion active users; \'Codex changed the game\'
OpenAI said active users of its AI models have surpassed 1 billion and the number of companies adopting them has topped 2 million. The figures were disclosed by Chief Financial Officer Sarah Fryer (사라 프라이어) in a company blog post. OpenAI said usage intensity is rising and internal agent-style work via Codex accounts for 99.8 percent of weekly output tokens. It also detailed operational optimisations and cost reductions alongside a recent price cut for GPT-5.6 models.
AI & Enterprise
OpenAI says GPT-5.6 Sol optimises GPU efficiency itself, cuts inference costs 20 percent
OpenAI says it used its latest AI model, GPT-5.6, to optimise its own AI infrastructure, lowering inference costs and improving throughput. It said improvements to the inference engine and GPU use cut end-to-end service costs by 20 percent for the GPT-5.6 Sol model and raised token generation efficiency by 15 percent. OpenAI also optimised KV cache handling, GPU workload distribution and agent operations, limiting tool output and improving caching.
AI & Enterprise
Cerebras compresses 163 seconds into 5 seconds, says GPU era is over
AI chip designer Cerebras put the 1 trillion-parameter open-weight model Kimi K2.6 into its enterprise inference service and achieved 981 tokens per second, a pace it says is the world’s fastest. It also cut the time to complete 500 output tokens from a 10,000-token input to 5.6 seconds, versus 163.7 seconds on the official Kimi endpoint. The company is pursuing an IPO and reported 2025 revenue of $510 million and net profit of $238 million.