Monday, August 17, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeCompute & CloudReport
Compute & Cloud · Report

French startup Kog is developing GPU optimization software to achieve 30x faster AI inference efficiency from existing Nvidia hardware.

Software-layer inference optimization is emerging as competitive alternative; Kog's approach unlocks productivity from sunk GPU capital and challenges Nvidia's vertical stack dominance.
Trade pressSlicast · August 15, 2026 · US · Source: Google News
importance 68

The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there's significantly more power to be squeezed out of conventional GPUs.

The startup garnered attention in May with a tech preview aimed at proving that "extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own" — such as the AMD MI300X and Nvidia H200 GPUs it used for its demo.

While some were disappointed to hear this didn't extend to laptops, others saw the potential. With inference speed and costs now being a critical bottleneck, Kog's promise to unlock new capabilities on existing hardware through software optimization attracted substantial interest. "We had 200 tangible business leads," CEO Gaël Delalleau told TechCrunch.

Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes wait hours for results. Anthropic itself understands that speed has value, charging a price multiple for Claude's Fast Mode.

Kog is targeting customers put off by those delays, typically those relying on AI workflows for professional tasks. The startup also counts design partners that generate games and apps with prompts, for whom faster inference through the Kog Inference Engine (KIE) would mean increased revenue, Delalleau said.

The company recognizes the market is not quite mature yet. From early customer conversations, Kog learned that prospective customers aren't prepared to fine-tune small models. "And that's why since the launch, we've been fully focused on accelerating the development of larger models to meet the demand we've seen."

This leaves Kog with a substantial challenge to deliver on its promise of "30x faster LLM inference." Its demo showed an impressive 3,000 per-request tokens per second — but used a purpose-built small model with only some 2 billion parameters, the now open-sourced Laneformer 2B.

Contradicting skeptics, Delalleau is confident the same approach can work equally well with larger LLMs, whose size can challenge inference chips. "GPUs have a bright future," he said. For Kog's CEO, the idea that they're ill-suited for decoding has become a misconception; newer GPUs have increasing memory bandwidth that only begs to be unlocked.

Kog isn't alone in thinking software optimization can help GPUs perform beyond their specifications. ZML, also from France, released hardware-agnostic software that bypasses Nvidia's CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper focus on GPU acceleration.

Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, bears no connection to Kog — except through his former co-founder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog's seed round. But the startup's deep technical focus stems from his unique background.

Having studied solid-state physics at École Polytechnique, Delalleau worked in offensive cybersecurity — white hat hacking. According to him, this shaped the mindset he's encouraging his team to adopt. On the science side, "there's this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them."

As for hacking, the four-time finalist at DEFCON's CTF tournament learned to "reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn't necessarily designed."

The downside of this approach is that it is hands-on and time-consuming. "For every new GPU, we'll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware." With a team of 11 people, this limits how many chips Kog can support in the near term.

Longer term, Kog hopes to feed its methodology into agent-based pipelines that will enable support for more chips and models. As Europe seeks to build capability in both areas, this could provide sovereignty tailwinds for the startup, already supported by Scaleway and backed by France's Bpifrance and French Tech 2030's program.

For now, though, Kog must prove its approach works on LLMs. This will also prove key to securing further funding. "Once we've implemented our first major model at 10x speed, which I think will be in September, we'll be able to start demonstrating customer traction and from there, raise our Series A," Delalleau said.

Read the original
French startup Kog is developing GPU… · Slicast