AI NewsKog is going deeper to squeeze more inference out of GPUs
Kog is going deeper to squeeze more inference out of GPUs
9:23 PM IST · August 14, 2026

The race for fasterAI inferenceis on, and markets gave Cerebras and its purpose-built chipsa warm welcomein its IPO debut in May. But French startupKogis betting that thereâs a lot more power to be squeezed out of conventional GPUs. The startuphit the front page of Hacker Newsin May with atech previewaimed at proving that âextremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already ownâ â such as the AMD MI300X and Nvidia H200 GPUs it used for its demo. Some were disappointed to hear this didnât extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kogâs promise to unlock new capabilities on existing hardware with software optimization attractedmore than onlookers. âWe had 200 tangible business leads,â CEO GaĂ«l Delalleau told TechCrunch. Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and chargesa price multiplefor Claudeâs Fast Mode. Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said. The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers arenât prepared to fine-tune small models. âAnd thatâs why since the launch, weâve been fully focused on accelerating the development of larger models to meet the demand weâve seen.â This leaves Kog with a huge leap to make to deliver on its promise of â30x faster LLM inference.â Its demo showed an impressive 3,000 per-request tokens per second (TPS) â but with a purpose-built small model with only some 2 billion parameters, thenow open sourcedLaneformer 2B. Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. âGPUs have a bright future,â he said. For Kogâs CEO, the idea that they arenât well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked. Kog isnât alone in thinking that software optimization can help GPUs do more than it says on the box.ZML, also from France, released hardware-agnostic software that bypasses Nvidiaâs CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University labHazy Research, with an even deeper-level focus on GPU acceleration. Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alumStribe, has nothing to do with his new one â other than his former co-founder turned VC Kamel Zeroual, whose firmVarsity VCco-led Kogâs seed round. But the startupâs deep-level focus stems from his unique background. Having studied solid-state physics at Franceâs Ăcole Polytechnique, he went on to work in offensive cybersecurity â also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, âthereâs this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.â As for hacking, the four-time finalist atDEFCONâs CTF tournamentsaid it taught him âto reverse-engineer things at a very low level â down to assembly language and binary code â to understand how it works, and to try to use it to achieve a goal for which it wasnât necessarily designed.â The downside of this approach is that it is very hands-on and time-consuming. âFor every new GPU, weâll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.â With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future. In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is alreadysupported by Scalewayand backed by FranceâsBpifranceandFrench Tech 2030âs program. For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. âOnce weâve implemented our first major model at 10x speed, which I think will be in September, weâll be able to start demonstrating customer traction and from there, raise our Series A,â Delalleau said.
read more