Gradium
Gradium is a voice API platform providing text-to-speech, speech-to-text, live translation, and voice cloning for production voice agents and automation systems. The platform prioritizes low latency (216ms time-to-first-audio on TTS) and accuracy on edge cases like phone numbers and email addresses without requiring text preprocessing. It integrates with Pipecat voice-agent frameworks and LiveKit real-time communication infrastructure, and offers on-device TTS for offline deployment. Pricing is consumption-based at $0.03 USD per 1,000 characters. Agencies use Gradium to build voice automation systems, transcribe client calls, deploy multilingual voice experiences, and clone client voices for brand consistency.
Gradium is a voice API platform providing text-to-speech, priced at $13/month on the XS plan, integrating with Pipecat, LiveKit, Coval, and Hugging Face. InnovaAI scores it 4.9/10 for agency adoption, best for Developer, Project Manager, and Strategist roles handling 5+ client meetings per week.
Agency Audit
Gradium provides production-grade voice APIs for text-to-speech, speech-to-text, live translation, and voice cloning with sub-250ms latency. Agencies building voice agents or automating voice-heavy workflows benefit most: teams integrating with Pipecat or LiveKit can deploy voice capabilities without managing separate transcription and synthesis vendors. The platform handles hard cases like phone numbers and email addresses without preprocessing, reducing QA cycles for voice-agent teams.
3recommended
36/mo
$2,687/mo
Moderate
Illustrative scenario. Not a guarantee. Net capacity is the value of reclaimed time at $75/hr, less the lowest verified paid base plan (flat plan cost is shared). Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- Developer handling voice agent development and deployment
- Project Manager handling client call transcription and note capture
- Strategist handling voice automation prototyping and pitch
- Your agency does not build voice applications or voice automation systems and your workflows center on text-based content, design, or campaign management.
- Your team lacks in-house engineering capacity to integrate APIs and you depend entirely on no-code platforms or third-party integrations for all tooling.
- Your transcription and voice synthesis needs are episodic or low-volume (fewer than 10 voice projects per year), making the engineering investment disproportionate to the payoff.
Internal Adoption Path
$13/mo
$13/mo flat plan
36 hr/mo
3 seats × 12 hr each
$2,700/mo
modeled at $75/hr labor rate
$2,687/mo
value − subscription cost
In this model, 3 seats reclaim 36 hours of team time each month. Valued at $75/hr that is $2,700/mo, and after the $13/mo subscription it leaves $2,687/mo of capacity for billable client work.
Illustrative scenario. Not a guarantee. Uses the lowest verified paid base plan. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Gradium
Text-to-speech with sub-250ms latency
Converts text to natural speech with 216ms time-to-first-audio (P50) on production infrastructure. Developers building voice agents or IVR systems use this to deploy realistic conversational experiences without noticeable delay between user input and agent response.
Speech-to-text transcription
Transcribes audio to text in real time. Account Executives and Project Managers use this to capture client call content automatically without manual note-taking, freeing attention for active listening and relationship building during calls.
Live voice-to-voice translation
Translates spoken input to spoken output across languages in real time. Agencies managing multilingual client projects or international support automation use this to reduce translation bottlenecks and deploy voice agents across regions without language-specific engineering.
Voice cloning for brand consistency
Clones client voices for use in voice agents or automated systems. Project Managers deploying voice automation for clients use this to maintain brand voice without expensive voice talent sessions, compressing project timelines by 1-2 weeks.
On-device TTS for offline deployment
Runs text-to-speech locally on client infrastructure without cloud calls. Developers building voice systems for regulated or air-gapped environments use this to meet compliance requirements while maintaining low latency.
Gradbot single-prompt agent builder
Builds voice agents from a single text prompt without manual orchestration. Strategists and Project Managers use this to prototype voice automation concepts in hours instead of days, accelerating client pitch cycles and proof-of-concept validation.
What Makes Gradium Different
Unique advantages vs similar tools in this niche
TTS model passes 81.0% of hard-case evaluation set
vs ElevenLabs v3 Conversational at 65.4%Gradium TTS outperforms competitors on reading phone numbers, emails, and reference codes correctly.
TTFA P50 of 216 ms with 30 ms p75-p25 spread
vs Cartesia Sonic 3.6 at 454 ms median with 165 ms spreadGradium offers lower and more consistent latency for live voice calls.
No hidden text normalization or rewriting
vs Other TTS labs that rewrite text with LLM before synthesisGradium ensures the studio demo matches production output exactly.
Latest Updates
Recent releases and improvements for Gradium
New Gradium Text-to-Speech Model Is Now Live
New2026-08-31A new Gradium TTS model is now the default. It handles hard real-world cases like phone numbers, email addresses, and IBANs with no pre-processing. TTFA P50 is 216 ms on Coval, 170 ms faster than the previous model, with a 30 ms p75-p25 spread.
Gradium Extends Funding to $100 Million and Expands to Silicon Valley
2026-07-08Gradium extends its seed funding to $100 million, welcoming new investors including NVIDIA, and opens a San Francisco Bay Area office to scale its real-time voice AI.
Value Equation
Outcome-likelihood-time-effort assessment for Gradium
Limited agency channel
Gradium scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.
Contact GradiumPricing
Gradium platform cost to your agency
XS: $13/mo
Free
- 45k included credits
- 1hr TTS audio
- 4hrs STT audio
- 3hrs STT Translation audio
XS
- 225k included credits
- 5hrs TTS audio
- 21hrs STT audio
- 16hrs STT Translation audio
Enterprise
- Unlimited credits
- Unlimited TTS, STT, and S2S Translation audio
- Unlimited custom voices
- Unlimited pro voice clones
Add-ons
Optional extras priced on top of any main plan
No verified white-label program for Gradium: client-facing delivery runs under the platform's native branding.
Market Intelligence
Offer + scale economics for Gradium
Limited agency channel
Gradium scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.
Contact GradiumInvestment Decision Framework
Strategic vetting analysis for Gradium
Situational Fit
Fit depends on your client mix
Buy If
4You operate a customer support automation practice and need sub-300ms voice latency to deploy realistic IVR or voice-bot prototypes without noticeable delay, which your current stack cannot guarantee.
Your development team builds voice agents or conversational AI products for clients and currently stitches together separate TTS, STT, and translation vendors, losing engineering time to integration overhead.
Your Project Managers or Strategists spend 3+ hours per week manually transcribing client calls or recording sessions because your current transcription tool requires third-party bot participation that clients reject.
Your team clones client voices for brand consistency in voice-agent deployments and currently relies on manual voice recording sessions that delay project timelines by 1-2 weeks per client.
Skip If
4Your agency does not build voice applications or voice automation systems and your workflows center on text-based content, design, or campaign management.
Your team lacks in-house engineering capacity to integrate APIs and you depend entirely on no-code platforms or third-party integrations for all tooling.
Your transcription and voice synthesis needs are episodic or low-volume (fewer than 10 voice projects per year), making the engineering investment disproportionate to the payoff.
Your clients require on-premises or air-gapped deployment and you cannot use cloud-based APIs, since Gradium does not publish self-hosted licensing options.
Bottom Line
Gradium provides production-grade voice APIs for text-to-speech, speech-to-text, live translation, and voice cloning with sub-250ms latency. Agencies building voice agents or automating voice-heavy workflows benefit most: teams integrating with Pipecat or LiveKit can deploy voice capabilities without managing separate transcription and synthesis vendors. The platform handles hard cases like phone numbers and email addresses without preprocessing, reducing QA cycles for voice-agent teams.
Reality Check
Gradium is a developer-first API platform, not a no-code tool. Your team needs engineering bandwidth to integrate it into existing workflows or agent infrastructure. Adoption ROI concentrates on agencies that run 5+ voice-agent projects annually or maintain ongoing voice automation systems.
Moderate effort: standard configuration with some customization needed
Academy for Gradium
Work through it in order: the course for this service first, then the modules behind it.
No Academy modules are published for this service yet. Browse the full Academy
Why this category matters
The commercial case before the tooling.
Core concepts
The mental model you need to price and scope the work.
- Post-Deployment Labor FloorConcept
Every voice agent deployment leaves a labor floor: the calls, escalations, and corrections that still need a person. The framework asks agencies to measure that floor before pricing a retainer, because the floor, not the license fee, decides whether the account is profitable. Start with real call samples: count how many calls the agent resolves end-to-end, how many escalate, and how many need a human to fix a booking or a misread intent. Trillet's identity verification and live-system actions raise the automation ceiling in regulated work, but a wrong payment action still lands on someone's desk. Ruby and Abby keep humans in the loop by design, so their floor is visible in the invoice; white-label platforms hide it until month two. Forrester's finding that 83% of B2C marketers already use AI agents means clients compare your offer against a baseline, so quote the floor explicitly or absorb it silently.
- Escalation Accuracy CeilingConcept
Escalation accuracy is the share of calls a voice agent routes to a human at the right moment, neither too early nor too late. It sets the ceiling on what an agency can charge, because every misrouted call becomes a client-visible failure that erodes trust faster than any latency or voice-quality issue. A 92% containment rate sounds strong until the 8% that should have escalated includes a billing dispute or a clinical question. Trillet verifies caller identity and executes actions in live systems with a full audit trail, which is the kind of control that makes escalation rules defensible in regulated accounts. Agencies should price a voice retainer only after sampling 50 to 100 real calls and measuring both false escalations (wasted human minutes) and missed escalations (client risk). The gap between those two numbers is the actual margin and the actual liability.
- Consent Surface MappingConcept
Consent Surface Mapping treats every jurisdiction, call-recording rule, and disclosure requirement as a boundary that shrinks or expands where an AI voice agent can actually run. Agencies that map the consent surface before scoping a retainer avoid the common failure of deploying a working agent into a state or vertical where recording without disclosure is illegal, forcing a rebuild after the client has already seen a demo. The framework has three layers: jurisdiction (two-party consent states, GDPR, TCPA), vertical (healthcare, legal, financial), and channel (inbound vs outbound, live vs voicemail). Trillet's identity verification and audit trail exist precisely because regulated industries require provable consent at each layer. A concrete example: an agency pitching a missed-call follow-up agent to a dental group must confirm HIPAA handling and state recording rules before quoting, or the first live call becomes a liability event rather than a lead recovery win.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- Voice Agent Rule: Price After Call Samples, Not After DemosEvaluation Rule
Collect at least 50 real recorded calls from the client's own phone line, run them through the candidate platform, and price the retainer only from measured containment, escalation accuracy, and per-minute usage cost.
- When Call Volume Is Under 200 a Month, Fix the Phone Process Before Buying a Voice AgentEvaluation Rule
Measure missed-call revenue and handoff failure rate first, and only deploy a voice agent when the recovered value per month exceeds the platform fee plus the labor hours the client must still staff.
- AI Voice Agent Decision: White-Label Platform vs Single-Client BuildDecision Framework
IF an agency expects to run voice agents for three or more client accounts within two quarters, THEN a white-label platform (Synthflow, ConvoCore, Autocalls, Trillet) amortizes setup across retainers and keeps the brand in the agency's name. IF the agency has one anchor client with a narrow call flow and no resale ambition, THEN a single-client build on conversational infrastructure (Vapi, Retell AI, LiveKit) avoids platform margin and gives full control of latency and escalation rules.
- The Demo-Call Trap: Why AI Voice Agent Pilots Stall Before Retainer RenewalFailure Pattern
- The Minutes-Only Trap: Why AI Voice Agent Retainers Collapse When Nobody Owns the Escalation PathFailure Pattern
Delivery system
Blueprints and procedures for running it as a service.
- Missed-Call Recovery Voice Agent Offer (10-14 days)Implementation Blueprint
A productized deployment that puts an AI voice agent on the client's inbound line to answer, qualify, and book calls that currently ring out, with escalation rules written into the flow. Priced only after real call samples, latency, and consent requirements are measured.
- Call Sample Audit Before Retainer Pricing (Onboarding)Operating Procedure
- Escalation Boundary Mapping (Onboarding)Operating Procedure
- Missed-Call Recovery Handoff (Handoff)Operating Procedure
13 modules selected for Gradium
Frequently Asked Questions
Answers about pricing, setup, implementation
Gradium provides APIs for text-to-speech, speech-to-text, live translation, and voice cloning optimized for production voice agents. It handles edge cases like phone numbers and email addresses without preprocessing and integrates with Pipecat and LiveKit frameworks. Agencies use it to build voice automation systems, transcribe calls, and deploy multilingual voice experiences.
Gradium offers 3 pricing tiers, at $13/mo (XS).
Developers and engineers benefit most by integrating Gradium into voice-agent systems and reducing vendor stitching overhead. Project Managers save time on call transcription and voice-cloning workflows. Strategists and Founders accelerate voice-automation pitch cycles using Gradbot. Account Executives reduce manual note-taking during client calls by capturing transcripts automatically.
For a developer integrating Gradium into a voice-agent project, expect 4-6 hours saved per project cycle by eliminating separate TTS and STT vendor management. For a Project Manager transcribing 2-3 client calls per week, Gradium reclaims 2-3 hours weekly by automating transcription. Savings scale with project volume and call frequency.
Yes. Gradium is a developer-first API platform. Your team must integrate it into existing systems, voice-agent frameworks, or call-recording infrastructure. Gradbot offers a no-code entry point for prototyping, but production deployments require API integration and testing.
Proof-of-concept integration typically takes 1-2 weeks for a developer to test Gradium APIs and validate latency in your infrastructure. Full production rollout depends on your existing tech stack and whether you are building new voice systems or retrofitting existing ones. Agencies using Pipecat or LiveKit see faster integration.
Gradium offers on-device TTS for offline deployment on client infrastructure. Cloud-based APIs (STT, translation, voice cloning) require internet connectivity. If your projects require fully air-gapped systems, on-device TTS covers text-to-speech; transcription and translation remain cloud-dependent.
Gradium does not publish a data retention or deletion policy in publicly available documentation. Contact Gradium directly to clarify data handling, retention periods, and deletion procedures for transcripts, voice clones, and call recordings after account closure.