Sakana AI Launches Fugu Max and Fugu Ultra v2
Sakana AI launched Fugu Max and Fugu Ultra v2 via OpenAI-compatible API on September 10, 2026, cutting costs by 2x to 6x.

Stock photo for illustration only, not from the actual event
- Sakana AI launched Fugu Max and Fugu Ultra v2 available today via API
- Fugu Max optimizes routing to reduce operational costs by 2x to 6x
- Fugu Ultra v2 targets complex reasoning and full-stack software development
Sakana AI officially announced the launch of its new artificial intelligence models, Fugu Max and Fugu Ultra v2, on September 10, 2026. Both models are accessible today as hosted APIs that are compatible with OpenAI. Currently, there are no open weights available for self-hosting, and the service is not offered in the EU/EEA region.
Sakana's core philosophy centers on balancing real capability with actual cost. Sending a straightforward data lookup to a massive multi-trillion-parameter model is inefficient. An optimal system selects the most economical machinery capable of resolving the task at hand.
The Sakana team describes this balance using the Pareto frontier, where improving quality increases expenses, and reducing costs compromises performance. Both Fugu Max and Fugu Ultra v2 share a single core orchestration architecture, differing only in their optimization targets.

Stock photo for illustration only, not from the actual event
This release continues the company's fast deployment cadence following Fugu's beta entry in April, general availability in June, and the addition of Fugu-Cyber and a Claude Code interface in July. According to the technical report, Fugu models function as independent language models that dynamically construct an agentic scaffold for each query.
The system builds directly upon two research papers from ICLR 2026. TRINITY employs a lightweight evolved coordinator to assign Thinker, Worker, or Verifier roles, while the Conductor is trained via reinforcement learning to discover natural-language coordination strategies.
"Real workloads are judged on capability and cost together."
Sakana AI
Fugu Max broadens the pool of models it can orchestrate by incorporating various open-weights and specialized models, including the NVIDIA Nemotron family through a collaboration with NVIDIA. Fugu Max routes each task to the leanest capable model, placing it within striking distance of elite models at two to six times lower cost based on internal SWEFish benchmarks.
Multi-agent orchestration represents a vital shift in optimizing AI workloads. By dynamically delegating sub-tasks to smaller or specialized models through an intelligent coordinator, organizations can drastically cut API expenditures without sacrificing output quality on complex tasks.
Meanwhile, Fugu Ultra v2 focuses squarely on complex reasoning, autonomous research, and full-stack software development, achieving its most significant performance leaps during sustained reasoning over visual and structured data.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment