Sakana
Sakana Fugu is an AI orchestration engine that distributes tasks across a pool of open-weights and specialized models, selecting the most cost-efficient model capable of handling each subtask rather than routing everything through a single frontier model. It ships two modes: Fugu Max, which prioritizes cost efficiency across the broadest available model pool, and Fugu Ultra v2, which targets peak performance on complex multi-step reasoning, autonomous research, and full-stack software development tasks. The platform exposes an OpenAI-compatible API endpoint, integrates with NVIDIA Nemotron models and the Claude Code environment, and is designed so the underlying model pool can be swapped without changing the calling application.
Sakana is an AI orchestration engine, integrating with NVIDIA Nemotron and Claude Code. InnovaAI scores it 4.1/10 for agency adoption, best for Software Engineer, Operations Director, and Strategist roles handling weekly client-facing work.
Agency Audit
Sakana Fugu is an AI orchestration platform that routes tasks across open and specialized models to optimize cost and capability trade-offs. It offers Fugu Max for cost-efficient outputs and Fugu Ultra v2 for maximum capability on complex reasoning tasks. For software engineering teams, research-heavy agencies, and AI development shops, Sakana reduces token spend by 2-6x on equivalent outputs while maintaining frontier-grade performance. The platform's swappable model pool avoids vendor lock-in and integrates with NVIDIA Nemotron and Claude Code workflows. Adoption makes sense for agencies running high-volume AI inference workloads where per-token costs compound across team seats.
4recommended
40/mo
No paid plan published
Moderate
Illustrative scenario. Not a guarantee. Net capacity needs a verified paid base plan, and none is published for this service, so it is not modeled. Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- Software Engineer handling LLM-assisted code generation and debugging
- Operations Director handling multi-step autonomous research
- Strategist handling AI vendor cost monitoring and consolidation
- Your agency's AI usage is ad-hoc and under 10 tasks per week per seat. Setup and monitoring overhead will exceed any token savings.
- Your team uses only closed-API models (GPT-4, Claude) and has no internal engineering capacity to manage API routing or monitor orchestration logs. Sakana requires technical oversight that non-technical roles cannot provide.
- Your workflows depend on real-time streaming responses or sub-100ms latency. Orchestration adds routing overhead that may exceed your latency budget for live client-facing features.
Internal Adoption Path
No paid plan published
40 hr/mo
4 seats × 10 hr each
$3,000/mo
modeled at $75/hr labor rate
No paid plan published
Illustrative scenario. Not a guarantee. No verified paid base plan is published for this service, so subscription cost and net capacity are not modeled. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Sakana
Dynamic Task Routing
Fugu analyzes each incoming task and routes it to the leanest model in the pool capable of solving it, reducing token spend on simpler subtasks. Engineering teams running mixed-complexity workloads benefit most from this automatic cost optimization.
Fugu Max Cost-Performance Mode
Fugu Max orchestrates the largest available pool of open-weights and specialized models to deliver frontier-grade outputs at reduced token cost. Operations leads tracking AI spend can use this mode as the default for high-volume, routine generation tasks.
Fugu Ultra v2 Peak Capability Mode
Fugu Ultra v2 targets maximum performance on complex, multi-step reasoning and autonomous research tasks without relying exclusively on a single closed frontier model. Strategists and researchers running deep analysis workflows can invoke this mode when output quality is the primary constraint.
OpenAI-Compatible API Endpoint
Fugu exposes a drop-in API endpoint compatible with the OpenAI interface, so engineering teams can substitute it into existing integrations without rewriting client code. This reduces the rollout friction for agencies already calling OpenAI endpoints in their internal tools.
Claude Code Interface
The Claude Code integration allows software development teams to use Fugu's orchestration layer directly within their coding environment, routing code generation and debugging subtasks across the model pool rather than sending all requests to a single model.
NVIDIA Nemotron Model Pool
Fugu integrates NVIDIA Nemotron open models into its orchestration pool, giving engineering teams access to additional specialized model capacity without managing separate API credentials for each provider.
What Makes Sakana Different
Unique advantages vs similar tools in this niche
Orchestrates a swappable pool of open and specialized models to expand the Pareto frontier
vs Single monolithic foundation modelsFugu Max achieves best overall score on six benchmarks while costing 40-60% less than Sonnet 5, GPT 5.6 Terra, and Kimi K3.
Delivers frontier performance without reliance on proprietary frontier models
vs Closed ecosystems like Fable 5 or GPT-6-AstraFugu Ultra v2 achieves top scores on five of eight benchmarks without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.
Latest Updates
Recent releases and improvements for Sakana
Fugu Max and Fugu Ultra v2 Release
New2026-09-11Fugu Max expands the Pareto Efficiency Frontier by orchestrating a larger pool of open and specialized models including NVIDIA Nemotron, delivering frontier-grade results at 40-60% lower cost than competitors. Fugu Ultra v2 pushes peak performance higher on complex multi-step reasoning, visual reasoning, and software engineering benchmarks.
Fugu-Cyber and Claude Code Interface
NewOrchestration specialized for real-world domain workflows including cybersecurity and coding environments.
Sakana Chat Update and NVIDIA Partnership
NewOrchestration scaled to daily consumer use in Sakana Chat, with integration of NVIDIA Nemotron open models begun.
Fugu General Availability and Fugu Ultra v1
NewFugu reached general availability with Fugu Ultra v1, demonstrating that an orchestration layer can match closed frontier models on hard benchmarks.
Fugu Beta Launch
NewProved multi-agent orchestration works as a unified foundation model.
Value Equation
Outcome-likelihood-time-effort assessment for Sakana
Limited agency channel
Sakana scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.
Contact SakanaPricing
Sakana platform cost to your agency
Add-ons
Optional extras priced on top of any main plan
No verified white-label program for Sakana: client-facing delivery runs under the platform's native branding.
Market Intelligence
Offer + scale economics for Sakana
Limited agency channel
Sakana scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.
Contact SakanaInvestment Decision Framework
Strategic vetting analysis for Sakana
Situational Fit
Fit depends on your client mix
Buy If
4Your software engineering team runs 50+ AI-assisted coding tasks per week using Claude or GPT-4, and your monthly token spend exceeds $800. Sakana's dynamic routing to cheaper models for routine refactoring or test generation could reduce that bill by 40-60% while maintaining output quality.
Your research or strategy team uses multi-step AI workflows (research synthesis, competitive analysis, report generation) that currently chain multiple API calls. Sakana's orchestration layer consolidates those calls into a single endpoint, cutting token overhead and reducing latency for Project Managers coordinating research sprints.
Your Founder or Operations lead needs visibility into AI cost-per-task across the agency. Sakana's Pareto frontier approach lets you tune cost vs. capability per workflow, giving you granular control over which tasks get Fugu Max (cheap) versus Fugu Ultra v2 (premium).
Your team is locked into a single model provider (OpenAI, Anthropic) and wants supply-chain resilience. Sakana's pool of open-weights and specialized models, including NVIDIA Nemotron, lets you swap model sources without rewriting prompts or workflows.
Skip If
4Your team uses only closed-API models (GPT-4, Claude) and has no internal engineering capacity to manage API routing or monitor orchestration logs. Sakana requires technical oversight that non-technical roles cannot provide.
Your agency's AI usage is ad-hoc and under 10 tasks per week per seat. Setup and monitoring overhead will exceed any token savings.
Your workflows depend on real-time streaming responses or sub-100ms latency. Orchestration adds routing overhead that may exceed your latency budget for live client-facing features.
Your agency operates in a regulated vertical (healthcare, finance) requiring single-vendor audit trails and compliance documentation. Sakana's multi-model pool complicates compliance audits and may not meet your vendor-lock requirements.
Bottom Line
Sakana Fugu is an AI orchestration platform that routes tasks across open and specialized models to optimize cost and capability trade-offs. It offers Fugu Max for cost-efficient outputs and Fugu Ultra v2 for maximum capability on complex reasoning tasks. For software engineering teams, research-heavy agencies, and AI development shops, Sakana reduces token spend by 2-6x on equivalent outputs while maintaining frontier-grade performance. The platform's swappable model pool avoids vendor lock-in and integrates with NVIDIA Nemotron and Claude Code workflows. Adoption makes sense for agencies running high-volume AI inference workloads where per-token costs compound across team seats.
Reality Check
Sakana requires engineering or technical PM oversight to configure routing rules and monitor orchestration performance. Teams without existing API-based AI workflows will need integration work. ROI is strongest for agencies already spending $500+ monthly on AI inference; smaller teams may not see payback within 6 months.
Moderate effort: standard configuration with some customization needed
Academy for Sakana
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Sakana Fugu Agency Implementation, Cost-Optimized AI Task Routing
Learn how to architect multi-model AI workflows that route tasks to the most cost-efficient model capable of solving each subtask, then package this capability as a productized service for clients. This course covers Fugu Max and Fugu Ultra v2 mode selection, API integration patterns, token cost tracking, and how to position dynamic task routing as a premium retainer offering.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Why this category matters
The commercial case before the tooling.
Core concepts
The mental model you need to price and scope the work.
- Wiring Over WidgetsConcept
The AI agent itself is a commodity, but the value for agencies lies in the integration layer: connecting a pre-built agent to a client's CRM, calendar, and review cycle. This framework shifts focus from selecting the 'best' agent to mastering the wiring process. For example, an agency using Vendasta's white-label AI receptionist for a local business must configure it to match the client's booking rules and follow-up cadence, turning a generic tool into a tailored service. As agentic AI adoption grows (77% of decision-makers now run agents in production), clients expect this customization. Agencies that treat agents as components and invest in repeatable wiring processes can charge retainers for ongoing optimization, rather than one-off setup fees.
- Wiring Over WidgetsConcept
The AI agent market sells finished workers, but the strategic value for agencies lies not in the agent itself, which is increasingly a commodity, but in the wiring that connects it to a specific client's CRM, calendar, and review cycle. This framework, 'Wiring Over Widgets,' argues that agencies that treat agents as components rather than products win. The agent is the widget; the wiring is the integration, customization, and ongoing optimization that turns a generic tool into a tailored solution. For example, a white-label platform like Vendasta provides AI employees, but the agency's role is to configure them for each local business's unique lead flow and follow-up process. This wiring is where retainer pricing originates, as it requires ongoing maintenance and adjustment. Recent research shows that 88% of B2B marketers face foundational gaps, meaning clients need help not just deploying agents, but ensuring their operations can support them. Agencies that master the wiring can charge a premium for the irreducible value they add.
- Integration MoatConcept
The Integration Moat framework holds that the durability of an AI agent engagement is determined by how deeply the agent is wired into a client's existing systems, not by the agent's underlying capability. Since the agent itself is increasingly a commodity, the switching cost for the client lives in the integrations: the CRM fields mapped, the calendar sync, the review-cycle triggers, and the exception-handling rules. Agencies that invest in this wiring create a moat that competitors offering generic agents cannot cross. For example, a white-label platform like Vendasta lets an agency deploy an AI receptionist for a local business, but the real value is in configuring it to the client's booking flow and follow-up cadence. With 77% of AI decision-makers now running agentic AI in production, clients expect this depth, and agencies that deliver it convert one-off projects into retainers.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- AI Agents Rule: Wire the Agent, Not the ProductEvaluation Rule
Treat the AI agent as a commodity component and focus your value on the integration into the client's specific workflows, systems, and review processes.
- AI Agents Rule: Wire the Agent, Not the ProductEvaluation Rule
Treat the AI agent as a commodity component and charge for the integration into the client's specific systems and workflows.
- The Productized Agent Trap: Why AI Agent Services Stall Without Client-Specific WiringFailure Pattern
- The Agent-as-Product Trap: Why AI Agent Services Stall Without Client-Specific WiringFailure Pattern
8 modules selected for Sakana
Frequently Asked Questions
Answers about pricing, setup, implementation
Sakana Fugu is an orchestration platform that routes tasks across a pool of open and specialized AI models, selecting the leanest model capable of handling each specific subtask. It offers two operating modes: Fugu Max for cost-efficient high-volume outputs and Fugu Ultra v2 for maximum capability on complex, multi-step reasoning tasks. It connects to NVIDIA Nemotron models and the Claude Code environment, and exposes an OpenAI-compatible API endpoint.
Sakana Fugu is priced on token consumption rather than per seat. Input tokens are billed at $2 per 1M input tokens. Output tokens are billed at $6 per 1M output tokens. There is no published flat monthly seat fee; total cost scales directly with your team's token volume.
Engineering and software development team members benefit most, using Fugu's Claude Code interface and dynamic routing to reduce the cost and latency of code generation workflows. Strategists and Researchers running autonomous multi-step research tasks get a capability lift from Fugu Ultra v2. Operations leads managing AI vendor costs can consolidate spend onto a single endpoint. Founders evaluating vendor lock-in risk benefit from the swappable model pool architecture.
A conservative estimate for an engineering team member running 10 or more hours of LLM-assisted coding or research tasks per week is 2 to 4 hours saved on prompt iteration and cost-monitoring overhead, because automatic routing removes the manual decision of which model to call for each task. This estimate is based on workflow compression, not vendor-reported figures, and will vary with actual token volume and task complexity.
For a team with an existing OpenAI API integration, switching the endpoint to Fugu is a low-effort change that an engineer can complete in a few hours. For teams building a new integration from scratch, expect one to three days of setup time to configure credentials, test routing behavior, and establish token-budget monitoring before handing off to non-technical team members.
The primary friction is cost unpredictability: because billing is purely usage-based, teams without token-volume estimates may see variable monthly costs until they establish usage baselines. A secondary friction is that non-technical roles cannot use Fugu directly without a front-end tool built on top of the API, so the immediate beneficiaries are engineering and operations roles rather than the full agency team.
Fugu's OpenAI-compatible API endpoint means any tool or internal script already calling the OpenAI API can be redirected to Fugu with a credential swap and a base-URL change. The Claude Code interface supports direct use within that coding environment. NVIDIA Nemotron models are available within the pool without requiring a separate NVIDIA API account.
Sakana does not publish a specific data-retention or offboarding policy in available documentation. Before committing to production use, agency Operations leads should request a written data-handling agreement from Sakana covering what happens to prompt and output data upon account closure, particularly for any client-adjacent workflows.