APAIIF 亞太人工智能產業總會APAIIFAI Knowledge
Latest AI Technology

Sakana Fugu Ultra v2 Officially Released: Multi-Agent Orchestration Achieves 74.3% DeepSWE, $30 Per Million Output Tokens, Challenging Single Frontier Model Dominance

September 12, 20261 Views
Sakana Fugu Ultra v2 Officially Released: Multi-Agent Orchestration Achieves 74.3% DeepSWE, $30 Per Million Output Tokens, Challenging Single Frontier Model Dominance
Sakana AI
多代理協調
Fugu Ultra v2
AI模型
DeepSWE

Sakana Fugu Ultra v2: The Multi-Agent Orchestration Revolution Redefining AI Competition

Introduction: A New Paradigm Beyond Single Models

On September 11, 2026, Japanese AI research company Sakana AI officially released Fugu Ultra v2. This model's arrival is not merely a technical iteration — it represents a fundamental challenge to the mainstream assumption that "bigger models equal greater capability" that has dominated the AI industry.

The core philosophy of Fugu Ultra v2 is that by intelligently orchestrating multiple specialized models, it is possible to achieve or even surpass the performance of top-tier models without relying on any single frontier model. This philosophy carries profound strategic significance in the AI competitive landscape of September 2026.

Technical Architecture: The Fusion of TRINITY and Conductor

Academic Foundation

Fugu Ultra v2's technical architecture is grounded in two important papers Sakana AI presented at ICLR 2026:

The TRINITY Paper:

  • Proposes a learned task decomposition framework
  • The system automatically identifies sub-task structures within complex tasks
  • Dynamically assigns different types of agents (Thinker, Worker, Verifier)

The Conductor Paper:

  • Designs a conducting mechanism for multi-agent coordination
  • Achieves cross-model task routing and result integration
  • Establishes inter-agent communication protocols and conflict resolution mechanisms

Agent Role Division

Fugu Ultra v2 employs three core agent roles:

Agent Type Responsibility Applicable Scenarios
Thinker High-level reasoning and strategic planning Complex problem analysis, research planning
Worker Specific task execution and code generation Programming, data processing, document generation
Verifier Result review and quality control Code testing, fact-checking, logical verification

This division of labor enables the system to maintain high performance across different task types while avoiding the performance compromises inherent in single-model approaches.

Swappable Model Pool

A key design decision in Fugu Ultra v2 is the adoption of a "swappable model pool" architecture:

  • The system does not rely on any single proprietary frontier model
  • Can flexibly integrate open-source and commercial models
  • Effectively mitigates risks from vendor lock-in, API revocations, and geopolitical service interruptions

Deliberately Excluded Models: Notably, the v2.0 release deliberately excludes Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra from its model pool. This decision is not a technical limitation but a strategic statement from Sakana AI: Fugu Ultra v2's performance derives from its orchestration architecture itself, not from dependence on specific flagship models.

Performance Benchmarks: Five Best-in-Class Results

Internal Benchmark Results

Sakana AI reports that Fugu Ultra v2 achieves best or joint-best scores on five of eight internal benchmarks:

Key Results:

  • DeepSWE: 74.3 (software engineering task benchmark)
  • Chartography: 48.3 (chart understanding and generation benchmark)
  • GDP.pdf: Best score (complex document understanding benchmark)
  • Toolathon: Best score (tool usage efficiency benchmark)
  • SWEFish: Best score (software engineering benchmark)

Performance Positioning

Fugu Ultra v2 is positioned as the "high-capability, high-quality" tier, complementing the cost-optimized Fugu Max in the same product line:

  • Fugu Max: Prioritizes cost-effectiveness, suitable for high-frequency, standardized tasks
  • Fugu Ultra v2: Prioritizes peak performance, suitable for complex reasoning, autonomous research, and full-stack software development

The Importance of Independent Verification

It should be noted that as of September 2026, the above performance data is self-reported by Sakana AI and has not yet been independently reproduced or placed on major third-party leaderboards like Artificial Analysis. This is important context when evaluating these figures.

Pricing and Availability

Pricing Structure

Fugu Ultra v2 offers transparent tiered pricing:

Standard Context (≤272K tokens):

  • Input: $5 per million tokens
  • Output: $30 per million tokens
  • Cached input: $0.50 per million tokens

Long Context (>272K tokens):

  • Input: $10 per million tokens
  • Output: $45 per million tokens
  • Cached input: $1.00 per million tokens

Technical Specifications

  • Context Window: 1 million tokens
  • Maximum Output Length: 128,000 tokens
  • API Compatibility: OpenAI-compatible API
  • Access Channels: Inception API, Baseten, OpenRouter

Regional Restrictions

Currently, the Fugu product line is not available in the European Union or European Economic Area, as Sakana AI is working to ensure compliance with GDPR and local regulations. This restriction has less impact on Asia-Pacific users but is an important consideration for European enterprise customers.

Strategic Significance for the AI Industry

Challenging the "Bigger is Better" Mainstream Assumption

The release of Fugu Ultra v2 represents an important challenge to the mainstream development path of the AI industry. For a long time, the primary path to improving AI performance has been increasing model parameter counts and training data scale. Sakana AI's research demonstrates that by intelligently orchestrating multiple medium-scale specialized models, it is possible to achieve or even surpass the performance of single super-large models on specific tasks.

Impact on Enterprise AI Procurement

For enterprise users, the emergence of Fugu Ultra v2 provides a new dimension of choice:

  1. Vendor Diversification: No longer dependent on a single AI provider
  2. Cost Controllability: Optimize costs through task routing
  3. Risk Distribution: Avoid business disruption from single vendor service outages

Asia-Pacific Adoption Prospects

Asia-Pacific enterprises, particularly in technologically mature markets like Singapore, Japan, and South Korea, demonstrate high acceptance of multi-agent orchestration architectures. As a Japanese AI company, Sakana AI has natural cultural and linguistic advantages in the Asia-Pacific market, potentially enabling faster adoption rates in the region.

Conclusion: The Era of Orchestrated Intelligence

The release of Sakana Fugu Ultra v2 marks the AI competition entering a new phase: shifting from pure model capability competition to orchestration architecture competition. In this new paradigm, the winner is not necessarily the company with the largest model, but the company that can most effectively coordinate and integrate multiple AI capabilities.

For practitioners and observers in the AI industry, Fugu Ultra v2 provides an important conceptual framework: on the path to pursuing AI capability, orchestrated intelligence may be more sustainable and flexible than pure scale expansion.

Competitive Context: Where Fugu Ultra v2 Fits in the AI Landscape

The Multi-Agent Orchestration Market

Fugu Ultra v2 enters a rapidly evolving market for multi-agent orchestration systems. Several other players are pursuing similar approaches:

  • GitHub Copilot HydraFusion: Microsoft's research preview using multi-model dynamic orchestration, claiming 67% cost reduction through Single/Cascade/Critique routing modes
  • OpenHands 1.0: Open-source autonomous coding agent achieving 72% SWE-bench accuracy
  • Cursor Router: Developer tool for automatic model selection based on task type, yielding 30-50% cost savings

Fugu Ultra v2 differentiates itself through its academic grounding (ICLR 2026 papers), its explicit exclusion of top frontier models to prove architectural merit, and its focus on peak performance rather than cost optimization.

Benchmark Comparison Context

To understand Fugu Ultra v2's 74.3% DeepSWE score in context:

System DeepSWE Score Architecture
Fugu Ultra v2 74.3% Multi-agent orchestration
OpenHands 1.0 ~72% Single autonomous agent
GPT-6 Astra Not disclosed Single frontier model
Claude Fable 5.1 Not disclosed Single frontier model

The comparison with OpenHands 1.0 is particularly instructive: both achieve similar performance on software engineering tasks, but through fundamentally different approaches. OpenHands uses a single powerful agent, while Fugu Ultra v2 orchestrates multiple specialized agents.

The Vendor Lock-in Problem and Fugu's Solution

One of the most significant strategic advantages of Fugu Ultra v2's architecture is its approach to vendor lock-in. The AI industry has seen several instances where:

  • API pricing changes dramatically (OpenAI's multiple pricing revisions)
  • Models are deprecated without warning (GPT-3.5 Turbo deprecation)
  • Geopolitical factors restrict access (export controls on advanced AI models)
  • Service outages affect dependent applications

By using a swappable model pool, Fugu Ultra v2 allows enterprises to maintain continuity even when individual model providers experience disruptions. This resilience is increasingly valuable as AI becomes critical business infrastructure.

Technical Deep Dive: The TRINITY Architecture

Task Decomposition in Practice

The TRINITY framework's approach to task decomposition works through several stages:

Stage 1: Task Analysis The orchestrator analyzes the incoming request to identify:

  • Task type (coding, reasoning, research, creative)
  • Complexity level (simple, moderate, complex)
  • Required capabilities (code execution, web search, document analysis)
  • Quality requirements (speed vs. accuracy tradeoff)

Stage 2: Agent Assignment Based on the analysis, the system assigns roles:

  • Complex reasoning tasks → Thinker agents with strong logical capabilities
  • Code generation → Worker agents optimized for programming
  • Result verification → Verifier agents with strong analytical skills

Stage 3: Coordination and Integration The Conductor framework manages:

  • Inter-agent communication protocols
  • Conflict resolution when agents disagree
  • Result synthesis and quality assessment
  • Escalation to human oversight when confidence is low

Why This Approach Outperforms Single Models on Specific Tasks

The multi-agent approach has inherent advantages for certain task types:

  1. Error Detection: When a Verifier agent independently checks a Worker agent's output, errors are caught before they propagate
  2. Specialization: Each agent can be optimized for its specific role, rather than compromising across all capabilities
  3. Parallel Processing: Multiple agents can work simultaneously on different aspects of a complex task
  4. Diverse Perspectives: Different agents may approach problems differently, reducing systematic biases

However, this approach also has limitations:

  • Higher latency for simple tasks (orchestration overhead)
  • More complex failure modes (agent coordination failures)
  • Higher cost per task compared to single-model approaches for straightforward requests

Enterprise Adoption Considerations

When to Choose Fugu Ultra v2

Fugu Ultra v2 is most appropriate for enterprises that:

  1. Have complex, multi-step workflows: The orchestration overhead is justified when tasks genuinely benefit from specialized agent roles
  2. Prioritize vendor independence: Organizations concerned about AI provider concentration risk
  3. Need peak performance on coding tasks: The 74.3% DeepSWE score makes it compelling for software development use cases
  4. Operate in Asia-Pacific: No EU restrictions, and Sakana AI's Japanese origins may provide cultural alignment advantages

When to Consider Alternatives

Fugu Ultra v2 may not be the best choice for:

  • Simple, single-turn queries where orchestration overhead adds unnecessary latency
  • Cost-sensitive applications where Fugu Max's lower pricing is more appropriate
  • Organizations requiring EU/EEA compliance (currently unavailable in those regions)
  • Applications requiring real-time responses where the orchestration latency is prohibitive

The Future of Multi-Agent Orchestration

Emerging Standards and Protocols

The multi-agent orchestration space is rapidly developing new standards:

  • Model Context Protocol (MCP): Emerging standard for agent-to-tool communication
  • Agent Communication Language (ACL): Protocols for inter-agent messaging
  • OpenAI's Swarm Framework: Open-source multi-agent coordination library

Fugu Ultra v2's OpenAI-compatible API positions it well to integrate with these emerging standards, potentially expanding its ecosystem compatibility over time.

The Path to Autonomous Research and Development

One of the most exciting potential applications of systems like Fugu Ultra v2 is autonomous scientific research. The combination of Thinker agents for hypothesis generation, Worker agents for experiment design and execution, and Verifier agents for result validation creates a framework that could accelerate research cycles significantly.

Early experiments in this direction are already underway, with AI systems being used to generate hypotheses, design experiments, and analyze results in fields ranging from drug discovery to materials science.

Conclusion: Orchestration as the New Frontier

Sakana Fugu Ultra v2 represents a significant step forward in the evolution of AI systems — from monolithic models to coordinated agent networks. Its 74.3% DeepSWE score, achieved without relying on the most powerful frontier models, demonstrates that architectural innovation can be as impactful as raw scale.

For the Asia-Pacific AI ecosystem, Fugu Ultra v2 offers a compelling alternative to dependence on US-based frontier model providers. As the region's AI adoption accelerates, systems that offer vendor independence, strong performance on coding and reasoning tasks, and flexible deployment options will find a receptive market.

The question is not whether multi-agent orchestration will become mainstream — it is which orchestration frameworks will emerge as the dominant standards. Sakana AI's academic rigor and transparent benchmarking approach position Fugu Ultra v2 as a serious contender in this emerging competition.

FAQ

Related Articles