Kog, a French startup, is pursuing software techniques to increase inference speed on standard datacenter GPUs, rather than relying on purpose-built AI chips. CEO Gaël Delalleau said the company demonstrated single-request decoding on AMD MI300X and Nvidia H200 hardware and received “200 tangible business leads.”
Demo results and public claims
The company’s technical preview showed 3,000 per-request tokens per second (TPS) using a small, open-sourced model, Laneformer 2B, which has roughly 2 billion parameters. Kog advertises a target of “30x faster LLM inference,” though the demo used a purpose-built small model and not a large production LLM.
Delalleau argued that modern GPUs have growing memory bandwidth that software can unlock, and that decoding on GPUs has been misunderstood. He contrasted Kog’s low-level GPU engineering approach with other software efforts, saying the startup is more akin to Hazy Research in its focus on deep GPU acceleration. He also noted ZML as another firm pursuing hardware-agnostic inference software.
Use cases, limits and engineering constraints
Based on early feedback, Kog expects initial customers to be software engineering teams and professional users who need lower-latency AI workflows. Delalleau cited examples of existing constraints such as long waits reported by some Claude users and the premium Anthropic charges for Claude’s Fast Mode as market signals that speed matters.
The company acknowledges significant engineering work remains to adapt its techniques to larger LLMs and to support many GPU models. With a team of 11, Kog said it typically spends several weeks or months reverse-engineering a new GPU to optimize performance, which limits short-term support to a small set of chips. The startup plans to feed its methodology into agent-based pipelines to scale support over time.
Kog is backed by Scaleway and supported by Bpifrance and the French Tech 2030 program; its seed round was co-led by Varsity VC. Delalleau stated the company aims to implement a first major model at 10x speed by September and to use that milestone to demonstrate customer traction ahead of a Series A raise.
Original source: TechCrunch AI