
Sakana AI's Fugu scores 72.4% on SWE-bench and undercuts GPT-5.5-Turbo on cost by orchestrating a swappable model pool
Sakana AI's Fugu scored 72.4% on SWE-bench Verified — surpassing GPT-5.5-Turbo at roughly one-third the inference cost — by training an orchestrator model to dynamically route tasks across a pool of specialist models, with no single-provider dependency baked in. For Japanese enterprises now facing export restrictions on Anthropic's Fable and Mythos models, the architecture offers a concrete path to frontier-level performance that a regulatory shift cannot revoke overnight.
Source: sakana.ai ↗
Fugu dynamically orchestrates the world's best models to tackle complex, multi-step tasks, accessible through a single model API.
Sakana AI
Why this matters
- → Frontier performance at one-third the cost of GPT-5.5-Turbo via dynamic model routing.
- → Eliminates single-vendor dependency; regulatory restrictions cannot instantly block access.
- → Orchestration architecture becomes operational necessity, not just technical optimization.
Orchestration over scale