Kyutai
Kyutai is an open-science AI research lab that publishes multimodal models for speech, vision, and language processing. Rather than operating as a managed API service, Kyutai releases trained models, source code, and training artifacts on GitHub and Hugging Face for teams to integrate directly into their own applications. Core models include Moshi (speech-native dialogue), Unmute (voice layer for text LLMs), Hibiki-Zero (real-time speech translation), Pocket TTS (CPU-native text-to-speech), and CASA (vision-language fusion). Agencies download these models, deploy them on their own infrastructure, and maintain them as dependencies in their product codebases.
Kyutai is an open-science AI research lab, integrating with GitHub, Hugging Face, Epic Games, and General Intuition. InnovaAI scores it 1.7/10 for agency adoption, best for Founder, Product Manager, and Engineer roles handling 5+ client meetings per week.
Agency Audit
Kyutai is an open-science AI lab that releases multimodal models for speech, vision, and language processing. Agencies building AI-native client deliverables or internal tools benefit most: product teams integrating Moshi for real-time voice dialogue, developers embedding Unmute to add voice to existing LLMs, and strategists prototyping vision-speech workflows via MoshiVis. The core value is access to production-ready, open-source models without vendor lock-in or per-API-call costs.
3recommended
36/mo
No paid plan published
High
Illustrative scenario. Not a guarantee. Net capacity needs a verified paid base plan, and none is published for this service, so it is not modeled. Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- Founder handling voice product prototyping
- Product Manager handling speech API integration and maintenance
- Engineer handling multimodal model evaluation
- Your agency does not build AI products or prototypes and instead delivers traditional design, strategy, or media services.
- Your team lacks in-house ML engineers or Python developers to integrate and maintain open-source models in production.
- You require SLA-backed support, managed infrastructure, and vendor accountability; Kyutai is a research lab, not a commercial platform.
Internal Adoption Path
No paid plan published
36 hr/mo
3 seats × 12 hr each
$2,700/mo
modeled at $75/hr labor rate
No paid plan published
Illustrative scenario. Not a guarantee. No verified paid base plan is published for this service, so subscription cost and net capacity are not modeled. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Kyutai
Moshi speech-native dialogue
Real-time voice conversation system that processes speech directly without text conversion, preserving emotion and non-verbal cues. Product teams use this to prototype voice-first client experiences and reduce latency in conversational AI deliverables.
Unmute text-to-speech integration
Adds low-latency streaming voice output to any existing LLM with a single integration point. Developers compress the workflow of retrofitting voice onto text-based tools from weeks of API wrangling to days of model integration.
Hibiki-Zero speech-to-speech translation
Real-time translation that preserves speaker identity and tone by operating on speech directly rather than converting to text. Strategists use this to scope multilingual voice-native client projects without per-minute translation API costs.
Pocket TTS CPU-native inference
100M-parameter multilingual text-to-speech model that runs faster than real-time on CPU, eliminating cloud dependency for voice synthesis. Teams deploying edge or offline-first tools avoid latency and licensing overhead of cloud TTS services.
MoshiVis and CASA multimodal fusion
Vision-language models that fuse visual inputs with speech or text, enabling prototypes of image-aware voice assistants. Product teams reduce iteration cycles on multimodal client deliverables by using open-source models instead of chaining proprietary APIs.
MuScriptor music-to-MIDI transcription
Lightweight model that transcribes music into multi-instrument MIDI notation. Audio-focused agencies use this to automate music analysis workflows or prototype music-aware AI tools for clients.
What Makes Kyutai Different
Unique advantages vs similar tools in this niche
Speech-native architecture processes audio directly instead of text round-tripping
vs Cascaded speech-to-text then text-to-speech pipelinesMoshi processes speech directly rather than converting to text and back, which means it has minimal latency and can understand emotions.
Open-science release of full model weights, code, and training datasets
vs Closed commercial AI APIs with usage-based licensingWe aim to build efficient models that integrate multiple input modalities, and to release the tools to use and understand them.
Lightweight TTS that runs on CPU faster than real-time
vs GPU-dependent text-to-speech modelsPocket TTS, our 100M-parameter multilingual TTS which runs on CPU faster than real-time, bringing voice ability to low resource devices with a single line of code.
Latest Updates
Recent releases and improvements for Kyutai
Pocket TTS training code released
New2026-08-25Training code for Pocket TTS, a 100M-parameter multilingual TTS model, has been released.
MuScriptor: Automatic Multi-instrument Transcription
New2026-07-10Release of MuScriptor, a lightweight open-source model for transcribing music into multi-instrument MIDI.
MIRA World Model
New2026-07-06Release of MIRA, a real-time multiplayer world model trained on 10k hours of Rocket League games.
Surflo: Consistent 3D surfaces from a global state
New2026-07-01Release of Surflo, a project focused on consistent 3D surface generation from a global state.
The FID Lottery
New2026-06-18Publication of The FID Lottery research.
Value Equation
Outcome-likelihood-time-effort assessment for Kyutai
Value math requires real pricing
The Value Equation (dream outcome × likelihood ÷ time × effort) feeds directly into ROI math. Kyutai has no published pricing, so we hold this section until real numbers are available.
Contact KyutaiPricing
Pricing data not yet available for Kyutai.
Reality Check
Kyutai is a research-first platform, not a managed SaaS. Adoption requires engineering effort to integrate models into your stack and maintain dependencies. Best ROI emerges only if your team actively builds AI products or prototypes; passive consumption of demos yields minimal operational lift.
High effort: requires technical configuration and team training
How This Accelerates White-Label Services
Who It's For
- ✓ai-product-teams-building-voice-interfaces
- ✓developers-integrating-speech-and-vision-models
- ✓research-and-applied-ai-labs
Acceleration Steps
- 1Schedule onboarding with the vendor
- 2Configure provide open-source speech-to-text and text-to-speech models
- 3Connect GitHub
- 4Launch your first client project
Academy for Kyutai
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Kyutai Agency Implementation, Building Voice-Native Client Products
Learn how to integrate Kyutai's open-source speech models into client deliverables, from deploying Moshi for real-time dialogue to adding voice layers with Unmute and scaling multilingual projects with Hibiki-Zero. This course teaches agencies how to architect voice-first solutions, manage model dependencies in production, and price voice-native services as retainers.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Core concepts
The mental model you need to price and scope the work.
- Inference Cost Pass-Through CeilingConcept
Inference Cost Pass-Through Ceiling is the point at which an agency can no longer absorb a model provider's price or latency change inside a fixed retainer, so the cost has to move to the client or the work has to shrink. The framework asks three questions per client engagement: what share of delivery cost is metered inference, how fast can that share be re-routed to a cheaper model, and what contract language lets you reprice. Forrester's 2027 predictions flag AI growth colliding with energy and infrastructure limits, which converts compute scarcity into API price movement on agency tools. A concrete case: an agency running document analysis on a frontier API can shift bulk classification to a smaller open-weight model served through Ollama or a gateway like Helicone, keeping the frontier model only for reasoning steps. That split is the ceiling defense.
- Provider Substitution WindowConcept
Provider Substitution Window is the interval during which an agency can move a client workload from one model provider to another without rewriting prompts, evals, or integration code. The window is widest at the orchestration layer and narrowest at the fine-tuned weights layer: a gateway swap takes hours, a retrained model takes a quarter. Agencies that measure this window per client account know exactly when they hold pricing leverage and when a vendor holds it. Forrester's 2027 predictions flag compute and energy constraints pushing API pricing upward, which turns a wide substitution window into a margin defense rather than an engineering nicety. A concrete case: an agency routing Claude and GPT traffic through a gateway such as Helicone or Portkey can shift a client's summarization workload in an afternoon when one provider raises rates, while a competitor with hardcoded SDK calls absorbs the increase on a fixed retainer.
- Margin Defense StackConcept
Margin Defense Stack treats AI infrastructure as a layered cost structure rather than a single line item. The bottom layer is raw compute and API tokens, the middle layer is routing and caching, and the top layer is the client-facing retainer price. Agencies that only negotiate the top layer absorb every shock from the layers beneath. Forrester's 2027 predictions flag that AI expansion is colliding with energy and infrastructure limits, which translates into API price increases for agency tools and compresses margins on AI-inclusive retainers. A concrete defense: route repeat prompts through a gateway such as Helicone or Portkey so cached responses cut token spend before it reaches the client invoice, and keep a local fallback like Ollama for privacy-sensitive work. When a client asks why the AI retainer costs what it does, the stack shows exactly which layer each dollar covers.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- When AI Margins Depend on Third-Party Compute, Price the Dependency Before You Sign the RetainerEvaluation Rule
Map every AI dependency in the delivery stack to a named provider, a fallback route, and a pass-through cost clause before quoting fixed-fee client work.
- AI Infrastructure Rule: Route Across Providers Before You Standardize on OneEvaluation Rule
Put a routing or gateway layer between your application and every model provider before any client deliverable depends on one vendor's endpoint.
- The Single-Provider Lock-In Trap in AI InfrastructureFailure Pattern
- The Token Bill Creep: Why AI Infrastructure Costs Outrun Agency RetainersFailure Pattern
8 modules selected for Kyutai
Frequently Asked Questions
Answers about pricing, setup, reliability
Kyutai is a Paris-based open-science AI lab that builds and releases multimodal models spanning speech, vision, and language. It publishes open-source code, trained models, and interactive demos including Moshi (speech-native dialogue), Unmute (voice for any LLM), Hibiki-Zero (real-time speech translation), and CASA (vision-language fusion). Agencies integrate these models into their own applications rather than consuming them as managed APIs.
Kyutai does not publish per-seat pricing. All models, code, and training artifacts are released open-source and free to download from GitHub and Hugging Face. Agencies incur only their own compute and hosting costs to run the models.
Product and AI teams building voice or multimodal client deliverables benefit most. Specifically: product managers scoping voice-native features, engineers integrating speech models into tools, strategists prototyping multilingual or vision-speech workflows, and founders evaluating foundation-model alternatives to proprietary APIs. Traditional design, strategy, and media roles derive minimal direct value.
Savings depend on your baseline workflow. If your team currently integrates third-party speech APIs (Deepgram, ElevenLabs, Google Cloud Speech), switching to open-source Kyutai models can save 8-12 hours per month on API management, cost reconciliation, and vendor lock-in mitigation. If you do not currently build voice products, Kyutai saves zero hours.
Yes. Kyutai is a research platform, not a managed SaaS. Your team must download models from GitHub or Hugging Face, integrate them into your codebase, manage dependencies, and handle inference infrastructure. Rollout typically requires 2-4 weeks of engineering work per model integration.
Kyutai publishes models and code on GitHub and Hugging Face. It has demonstrated integrations with Epic Games (for world-model applications), General Intuition, and Les Invincibles (for accessibility tools). Your team integrates Kyutai models directly into Python, JavaScript, or other frameworks; there is no pre-built connector ecosystem.