The contention for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a lukewarm welcome successful its IPO debut successful May. But French startup Kog is betting that there’s a batch much powerfulness to beryllium squeezed retired of accepted GPUs.
The startup hit the beforehand leafage of Hacker News successful May with a tech preview aimed astatine proving that “extremely accelerated single-request decoding is imaginable connected the modular datacenter GPUs enterprises already own” — specified arsenic the AMD MI300X and NVIDIA H200 GPUs it utilized for its demo.
Some were disappointed to perceive this didn’t widen to GPUs successful our laptops, but others saw the potential. With inference velocity and costs present being a captious bottleneck, Kog’s committedness to unlock caller capabilities connected existing hardware with bundle optimization attracted more than onlookers. “We had 200 tangible concern leads,” CEO Gaël Delalleau told TechCrunch.
Based connected aboriginal feedback, the solo laminitis expects bundle engineering to beryllium the archetypal usage case. Veteran Claude Code users are good alert that they sometimes person to hold hours to get results. Anthropic itself understands that velocity is worthy money, and charges a terms multiple for Claude’s Fast Mode.
Kog is hoping to people customers enactment disconnected by those delays, usually due to the fact that they trust connected AI workflows for nonrecreational tasks. But the startup besides has plan partners that fto users make games and apps with a prompt, and for whom a faster result acknowledgment to the Kog Inference Engine (KIE) would mean much revenue, Delalleau said.
The institution realizes this marketplace is not rather mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune tiny models. “And that’s wherefore since the launch, we’ve been afloat focused connected accelerating the improvement of larger models to conscionable the request we’ve seen.”
This leaves Kog with a immense leap to marque to present connected its committedness of “30x faster LLM inference.” Its demo showed an awesome 3,000 per-request tokens per 2nd (TPS) — but with a purpose-built tiny exemplary with lone immoderate 2 cardinal parameters, the now open-sourced Laneformer 2B.
Contradicting skeptics, Delalleau is assured the aforesaid attack tin enactment conscionable arsenic good with LLMs, whose size tin beryllium a situation for inference chips. “GPUs person a agleam future,” helium said. For Kog’s CEO, the thought that they aren’t good suited for decoding has go a misconception; newer GPUs person much and much representation bandwidth that lone begs to beryllium unlocked.
Kog isn’t unsocial successful reasoning that bundle optimization tin assistance GPUs bash much than it says connected the box. ZML, besides from France, released hardware-agnostic bundle that bypasses Nvidia’s CUDA to enactment accelerated inference crossed competing chips. But Delalleau said Kog is much akin to Stanford University laboratory Hazy Research, with an adjacent deeper-level absorption connected GPU acceleration.
Delalleau himself is not a researcher, and his archetypal startup, TechCrunch50 2009 alum Stribe, has thing to bash with his caller 1 — different than his erstwhile cofounder turned VC Kamel Zeroual, whose steadfast Varsity VC co-led Kog’s effect round. But the startup’s deep-level absorption stems from his unsocial background.
Having studied solid-state physics astatine France’s École Polytechnique, helium went connected to enactment successful violative cybersecurity — besides known arsenic achromatic chapeau hacking. According to Delalleau, this shaped the mindset helium is present encouraging his squad to adopt. On the subject side, “there’s this mindset of knowing the laws of physics, and the laws of the GPU successful bid to marque the astir of them.”
As for hacking, the four-time finalist astatine at DEF CON’s CTF tournament said it taught him “to reverse-engineer things astatine a precise debased level — down to assembly connection and binary codification — to recognize however it works, and to effort to usage it to execute a extremity for which it wasn’t needfully designed.”
The downside of this attack is that it is precise hands-on and time-consuming. “For each caller GPU, we’ll dedicate respective weeks oregon adjacent months, to truly excavation into the details and behaviour GPU engineering probe connected that hardware.” With a squad of 11 people, this puts a bounds to the fig of chips that Kog tin enactment with, astatine slightest for the foreseeable future.
In the longer run, Kog hopes to provender its methodology into agent-based pipelines that volition fto it enactment much chips and models. As Europe seeks to physique its ain capableness connected those 2 fronts, this could adhd sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program.
For now, though, Kog needs to beryllium to the satellite that its attack works connected LLMs. This volition besides beryllium cardinal to securing much funding. “Once we’ve implemented our archetypal large exemplary astatine 10x speed, which I deliberation volition beryllium successful September, we’ll beryllium capable to commencement demonstrating lawsuit traction and from there, rise our Series A,” Delalleau said.
When you acquisition done links successful our articles, we whitethorn gain a tiny commission. This doesn’t impact our editorial independence.















English (US) ·