| Mobile Web

Lightbits launches KV cache engine to boost GPU performance

Lightbits Labs has launched its AI inference software engine, Infera, aimed at improving inference performance and cost efficiency. The company says Infera expands KV cache data beyond limited high-bandwidth memory attached to GPUs and targets large language model operations with long context windows or many concurrent sessions. It manages KV cache across GPU memory, DRAM and NVMe storage using predictive prefetching and supplies context in small blocks.