Kog, a French startup, is working on software solutions to enhance the speed of AI inference on conventional GPUs. This approach contrasts with companies developing purpose-built AI chips, as Kog believes substantial performance gains can still be achieved from existing hardware through software optimization.
In a tech preview, Kog showcased its ability to achieve 3,000 tokens per second (TPS) for single-request decoding on standard data center GPUs, including the AMD MI300X and NVIDIA H200. This demonstration used a 2-billion parameter model called Laneformer 2B. The company reported receiving 200 business leads following the preview, indicating significant market interest in faster inference solutions.
The primary goal of Kog's technology is to alleviate the critical bottleneck of inference speed and associated costs in AI applications. Many professional AI users, such as those using Claude Code, experience long wait times for results. Kog aims to provide a faster outcome, which could translate to increased revenue for businesses relying on AI workflows for tasks like generating games and applications.
Initial market feedback revealed that prospective customers are not prepared to fine-tune smaller models. In response, Kog has shifted its focus to accelerating the development of its technology for larger language models. The company is confident that its approach can scale to larger LLMs, despite the challenges their size presents for inference chips, by leveraging the memory bandwidth of newer GPUs.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
French startup Kog is developing software to significantly speed up large language model (LLM) inference on standard data center GPUs like AMD MI300X and NVIDIA H200. This initiative aims to address the bottleneck of inference speed and cost, potentially unlocking new capabilities on current hardware for professional AI workflows.