
Sankalp drives Codex to a 232x faster QR kernel on B200s
Sankalp ran GPT-5.5 Codex through more than 1,500 submissions over 14 days on GPU Mode's batched compact-Householder QR problem, cutting B200 runtime from about 419,000 microseconds to 1,805 — 232x over baseline, and 12th of 183 entrants. The transferable piece is the harness, not the kernel: keep three to five parallel candidate families alive instead of one improvement path, profile through Modal and NVIDIA's Nsight Compute, and select on measured runtime rather than reviewer judgment. The loop was steered every two to three hours, and the number is self-reported and not independently reproduced.
Source: sankalp.bearblog.dev ↗
Agents yearn for tight feedback loops. They allow them to hill-climb to their heart's content.
Sankalp
Why this matters
- → Demonstrates 232x speedup on GPU kernels via agent-driven iteration, not architectural breakthroughs.
- → Tight feedback loops (1,500 submissions in 14 days) outperformed domain expertise; harness design matters more
- → Practical proof that auto-research on well-scoped numerical problems is accessible without PhD-level optimizat
Loop engineering beats expertise