Sakana's new models route your task to the cheapest engine that can handle it. You pay the same flat rate either way
Published: 9/12/2026 · Source: Sakana AI ↗
Sakana AI released two new versions of Fugu on 11 September 2026, and both are built on the same idea: instead of one enormous network, a pool of open-weight and specialised models, with a router that sends each task to the leanest member capable of solving it. "A system that deploys a multi-trillion-parameter model to execute a simple data lookup is not intelligent, but wasteful," the Japanese lab writes.
The interesting part is not in the announcement. It is in the price list.
The original Fugu, still on sale, bills the customer for whatever it picked: when one agent is active, the buyer pays "only the standard rate for that specific underlying model", and when several are active, a single rate based on the most expensive one involved. Clever routing on a cheap task shows up directly on the invoice.
The two new flagships do not work that way. Fugu Max is a flat $2 per million input tokens and $6 per million output. Fugu Ultra v2 is a flat $5 and $30, doubling to $10 and $45 above a 272,000-token context. The rate does not move with the engine that answered. Route a question to a small open model or to an expensive specialist, and the bill is identical.
That is a legitimate way to sell a service — flat pricing is predictable, and Sakana carries the risk if a task turns out to need the expensive machinery. But it inverts who benefits from the routing. Under the old scheme, every cheap decision landed in the customer's pocket. Under the new one, it lands in Sakana's margin, and the customer's gain is limited to the headline rate being lower than a frontier model's: the lab claims Fugu Max undercuts Sonnet 5, GPT 5.6 Terra and Kimi K3 on output pricing by 40 to 60 percent.
The second thing worth reading twice is what Fugu Ultra v2 does not contain. Sakana states plainly that Fable 5, Fable 5.1 and GPT-6-Astra are not in its agent pool — the results are meant to be reached without the strongest closed models on the market. The lab reports 48.3 on Chartography, a test of reasoning over charts and structured data, against 27.3 for Opus 5 and 29.5 for Fable 5, and 74.3 on DeepSWE. Every one of those numbers is the vendor's own; none has been independently reproduced.
What neither product has is a specification in the ordinary sense. There is no parameter count, no stated context window, no training-data description and no model card, because, as Sakana puts it, Fugu is an architecture rather than a fixed set of weights. The only public hint at the size of the window is that Ultra v2's price changes above 272,000 tokens. Buyers of these two models are not buying a network. They are buying a routing policy over other companies' networks — and, as of this release, paying a fixed fee for it.
Both models are live through Sakana's OpenAI-compatible API. A technical report accompanies the release. Their catalogue profiles on wujec.ai record what the maker discloses and, just as importantly, what it does not.