Thinking Machines
Thinking Machines provides Inkling-Small, a 276B-parameter Mixture-of-Experts model with 12B active parameters, optimized for cost-efficient multimodal inference. The model processes audio, images, and text natively, supports variable reasoning effort (minimal to xhigh), and maintains a 1M token context window. Agencies access it via Tinker Playground (web interface), Tinker API, or self-hosted deployment on Hugging Face. It supports fine-tuning for domain-specific tasks and agentic workflows (tool use, coding). Pricing is usage-based: $1.20 per 1M output tokens, one-quarter the cost of the larger Inkling model.
Thinking Machines is an AI infrastructure platform, integrating with Hugging Face, Tinker, Tinker Playground, and Artificial Analysis. InnovaAI scores it 3.8/10 for agency adoption, best for AI Product Developer, Data Scientist, and Engineer roles handling 5+ client meetings per week.
Agency Audit
Thinking Machines offers Inkling-Small, an open-weights Mixture-of-Experts model with 276B total parameters and 12B active, enabling agencies to run multimodal AI inference (audio, image, text) at one-quarter the compute cost of its larger sibling. The model supports variable reasoning effort, 1M token context windows, and agentic tool use, making it relevant for agencies building custom AI applications or deploying open-weights models internally. Best adoption fit: technical teams (AI product developers, data science consultancies, custom AI builders) who need cost-efficient inference for client work or internal automation. Non-technical agency teams (creative, account management, operations) derive minimal direct value unless your workflow already involves running inference tasks.
3recommended
36/mo
No paid plan published
High
Illustrative scenario. Not a guarantee. Net capacity needs a verified paid base plan, and none is published for this service, so it is not modeled. Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- AI Product Developer handling multimodal document analysis
- Data Scientist handling custom model fine-tuning
- Engineer handling inference cost optimization
- Your agency is primarily creative, account-driven, or operations-focused with no in-house ML or engineering capacity; Inkling-Small requires infrastructure setup and is not a consumer-grade chat tool.
- Your inference volume is below 5M output tokens per month; the per-token cost advantage does not offset the operational overhead of managing an open-weights deployment.
- Your team relies on closed-model APIs (OpenAI, Anthropic, Google) and has no appetite to manage model versioning, fine-tuning pipelines, or self-hosted inference infrastructure.
Internal Adoption Path
No paid plan published
36 hr/mo
3 seats × 12 hr each
$2,700/mo
modeled at $75/hr labor rate
No paid plan published
Illustrative scenario. Not a guarantee. No verified paid base plan is published for this service, so subscription cost and net capacity are not modeled. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Thinking Machines
Variable reasoning effort control
Adjust inference reasoning from minimal to xhigh per task, allowing AI product developers and data scientists to trade off compute cost and answer quality without retraining or switching models. Saves 20-40 percent of inference spend on routine tasks while preserving full reasoning depth for complex client deliverables.
Native multimodal reasoning
Process audio, images, and text in a single inference call without chaining separate transcription, vision, and language models. Reduces integration complexity for strategists and analysts building document-analysis or media-review workflows.
1M token context window
Ingest entire long-form documents, transcripts, or research corpora in a single request. Eliminates chunking and retrieval complexity for data scientists and engineers building knowledge-synthesis tools for client projects.
Fine-tuning on Tinker
Customize Inkling-Small for domain-specific tasks (e.g., legal contract analysis, technical documentation review) without retraining from scratch. Enables AI product developers to deliver specialized models to clients faster and at lower infrastructure cost.
Agentic tool use and coding
Execute function calls and write executable code natively within inference. Allows engineers to build autonomous workflows (API calls, data transformations, system commands) without external orchestration frameworks.
Open-weights deployment
Run Inkling-Small on your own infrastructure or via Hugging Face, avoiding vendor lock-in and data residency concerns. Critical for agencies serving regulated clients or operating in jurisdictions with data sovereignty requirements.
What Makes Thinking Machines Different
Unique advantages vs similar tools in this niche
Token-efficient reasoning with variable thinking effort
vs Larger models like Inkling that require more computeInkling-Small achieves comparable performance to Inkling at a quarter of the compute cost, with test-time compute curves above Inkling's on reasoning benchmarks.
Native multimodal audio and vision processing
vs Models that require separate audio and vision pipelinesUses an encoder-free architecture with dMel spectrograms and image patches, enabling joint processing of audio, images, and text.
Open-weights availability with fine-tuning support
vs Closed-weight models that restrict customizationFull weights are released on Hugging Face, and the model is available for fine-tuning on Tinker.
Latest Updates
Recent releases and improvements for Thinking Machines
Introducing Inkling-Small
New2026-07-30Release of Inkling-Small, an efficient open-weights Mixture-of-Experts model with 276B total parameters (12B active), featuring native reasoning over audio and images, variable thinking effort, and a context window of up to 1M tokens. Available for fine-tuning on Tinker and for chat on Tinker Playground.
Value Equation
Outcome-likelihood-time-effort assessment for Thinking Machines
Limited agency channel
Thinking Machines scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.
Contact Thinking MachinesPricing
Thinking Machines platform cost to your agency
Add-ons
Optional extras priced on top of any main plan
No verified white-label program for Thinking Machines: client-facing delivery runs under the platform's native branding.
Market Intelligence
Offer + scale economics for Thinking Machines
Limited agency channel
Thinking Machines scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.
Contact Thinking MachinesInvestment Decision Framework
Strategic vetting analysis for Thinking Machines
Situational Fit
Fit depends on your client mix
Buy If
5Your technical team (engineers, data scientists, AI product leads) runs 10M+ output tokens monthly on inference tasks and currently pays for closed-model APIs; Inkling-Small at $1.20 per 1M tokens versus $4.05 for Inkling reduces monthly inference spend by 70 percent.
Your AI product developers need to fine-tune models on Tinker for client deliverables and want to avoid vendor lock-in with proprietary closed-weights models.
Your strategists or data analysts process long-form documents (contracts, research, transcripts) and need native multimodal reasoning over audio and images without spinning up separate transcription or vision APIs.
Your team builds agentic workflows (tool use, coding tasks, forecasting) and needs to dial reasoning effort up or down per task to balance latency and accuracy without switching models.
You operate in a regulated or privacy-sensitive vertical and prefer open-weights models you can audit and deploy on your own infrastructure rather than sending data to third-party API providers.
Skip If
5Your agency is primarily creative, account-driven, or operations-focused with no in-house ML or engineering capacity; Inkling-Small requires infrastructure setup and is not a consumer-grade chat tool.
Your inference volume is below 5M output tokens per month; the per-token cost advantage does not offset the operational overhead of managing an open-weights deployment.
Your team relies on closed-model APIs (OpenAI, Anthropic, Google) and has no appetite to manage model versioning, fine-tuning pipelines, or self-hosted inference infrastructure.
You need real-time, always-on inference at sub-100ms latency for production customer-facing features; Inkling-Small's performance-compute trade-off is optimized for batch and offline tasks, not low-latency serving.
Your compliance or security posture forbids running third-party model code, even open-source; Thinking Machines does not publish SOC 2, HIPAA, or FedRAMP certifications.
Bottom Line
Thinking Machines offers Inkling-Small, an open-weights Mixture-of-Experts model with 276B total parameters and 12B active, enabling agencies to run multimodal AI inference (audio, image, text) at one-quarter the compute cost of its larger sibling. The model supports variable reasoning effort, 1M token context windows, and agentic tool use, making it relevant for agencies building custom AI applications or deploying open-weights models internally. Best adoption fit: technical teams (AI product developers, data science consultancies, custom AI builders) who need cost-efficient inference for client work or internal automation. Non-technical agency teams (creative, account management, operations) derive minimal direct value unless your workflow already involves running inference tasks.
Reality Check
Inkling-Small requires infrastructure familiarity; it is not a plug-and-play SaaS dashboard. Agencies without in-house ML engineers or DevOps capacity will face setup friction. Payback period depends entirely on inference volume: teams running fewer than 10M output tokens monthly will not recoup seat costs.
High effort: requires technical configuration and team training
Academy for Thinking Machines
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Thinking Machines Agency Implementation, Cost-Efficient Multimodal AI Delivery
Learn how to build profitable retainer services around Inkling-Small's 276B-parameter model, leveraging variable reasoning effort and native multimodal processing to reduce inference costs by 20-40 percent while maintaining quality for client deliverables. Master fine-tuning workflows, agentic tool use, and pricing strategies that turn token efficiency into recurring revenue.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Why this category matters
The commercial case before the tooling.
Core concepts
The mental model you need to price and scope the work.
- Multi-Model Margin ShieldConcept
Agencies integrating AI into client solutions face a hidden margin killer: lock-in to a single model provider. When one vendor raises prices or shifts capabilities, project feasibility and retainer margins erode overnight. The Multi-Model Margin Shield framework treats provider diversity as a financial hedge, not just a technical preference. By routing requests through an orchestration layer that can switch between Anthropic's Claude, OpenAI's GPT, and Google's Vertex AI based on cost and latency, agencies protect delivery margins and negotiate from strength. This approach also guards against capability shifts, such as when a model's safety guardrails change mid-project. For example, a recent study found GPT-6 Astra blocks 99.99% of direct prompt injections but fails 8.5% of hidden ones, while Claude Opus 5 performs differently, underscoring why redundancy matters for client-facing agents.
- Provider Substitution WindowConcept
Provider Substitution Window is the measure of how cheaply an agency can move a client workload from one model provider to another, and it sets the ceiling on what any single vendor can charge before the account walks. The window is widest when prompts, evals, and routing live in an abstraction layer rather than inside a provider SDK, and narrowest when fine-tunes, cached embeddings, and agent memory are tied to one endpoint. For agencies on retainer, window width is a margin instrument: a delivery team that can swap endpoints in an afternoon negotiates from a different position than one facing a rewrite. The window also has a security edge. Anthropic's 150-page misuse report documents eight months of Claude abuse, including 151 million exchanges logged by Alibaba's Qwen team, which is exactly the kind of finding enterprise clients raise in procurement reviews. An agency that can answer with a documented swap path keeps the account.
- Orchestration Layer Lock-InConcept
Agencies integrating frontier models like Anthropic's Claude or OpenAI's GPT-5.6 into client solutions face a hidden risk: direct API dependency. Pricing changes, capability shifts, or outages at a single provider can erode project margins overnight. The framework of Orchestration Layer Lock-In argues that agencies should treat the model provider as a commodity and invest in a multi-model orchestration layer that abstracts routing, fallbacks, and cost management. This layer, exemplified by gateways like Helicone or OpenRouter, lets agencies switch between Claude, GPT, or others without rewriting client code. For instance, when Meta's ad AI altered approved creative post-launch, agencies relying on a single platform had no recourse; an orchestration layer would have enabled rapid failover to a safer model. By decoupling delivery from any one vendor, agencies protect margins and maintain negotiating power.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- AI Infrastructure Rule: When Lock-In Risk Rises, Route Through an Abstraction LayerEvaluation Rule
Before scaling any AI-powered client deliverable, route requests through a gateway or orchestration layer that supports multiple model providers.
- AI Infrastructure Rule: When Agent Workloads Scale, Gate Every Model Call Through an Observability ProxyEvaluation Rule
Route every model request through an observability and gateway layer before scaling any agent workload to more than one client.
- Multi-Model Orchestration Layer vs Single-Provider DependencyDecision Framework
IF your agency integrates frontier models into client deliverables and cannot absorb sudden pricing or capability shifts, THEN build a multi-model orchestration layer that routes requests across providers. IF your client work is low-volume, prototype-stage, or tightly coupled to one model's unique behavior, THEN a single-provider dependency is acceptable until scale justifies abstraction.
- The Single-Provider Lock-In Trap in AI InfrastructureFailure Pattern
- The Cost-Latency Blind Spot in AI InfrastructureFailure Pattern
Delivery system
Blueprints and procedures for running it as a service.
- Multi-Model AI Gateway & Observability Sprint (7-14 days)Implementation Blueprint
A structured engagement to design and deploy a vendor-neutral AI infrastructure layer for client applications, reducing lock-in risk and providing cost, latency, and reliability controls.
- Multi-Provider Model Orchestration Review (QA)Operating Procedure
- Provider Lock-In Risk Assessment (Onboarding)Operating Procedure
- AI Cost Governance Review (Retention)Operating Procedure
13 modules selected for Thinking Machines
Frequently Asked Questions
Answers about pricing, setup, implementation
Thinking Machines offers a free plan; paid pricing is not published publicly.
Inkling-Small costs $1.20 per 1M output tokens. The larger Inkling model costs $4.05 per 1M output tokens. Both are usage-based; no monthly seat fees or minimum commitments. An agency running 50M tokens monthly on Inkling-Small would spend $60; the same volume on Inkling would cost $202.50.
AI product developers and engineers gain the most direct value, using Inkling-Small to build custom inference pipelines and fine-tuned models for client deliverables. Data scientists and strategists benefit when analyzing long-form documents or building forecasting workflows. Operations and founders see indirect ROI through reduced inference spend if your team already runs 10M+ tokens monthly on closed-model APIs.
Hours saved depend entirely on inference volume and current tooling. An engineer currently chaining three separate APIs (transcription, vision, language model) for multimodal tasks saves 3-5 hours per week on integration and debugging. A data scientist running 50M+ tokens monthly on closed-model APIs saves 2-4 hours per month on cost optimization and vendor management. Non-technical roles see negligible time savings.
Inkling-Small is available via Tinker Playground (web chat), Tinker API, and Hugging Face. It integrates with Scale AI, Artificial Analysis, and forecasting platforms (Forecasting Research Institute, ProphetArena). If your team uses Python, JavaScript, or REST APIs, integration is straightforward. It does not natively connect to Slack, Notion, or project management tools; custom middleware is required.
Thinking Machines does not publish a data retention or deletion policy in public documentation. If you self-host on Hugging Face or your own infrastructure, you retain full control. If you use Tinker Playground or API, contact Thinking Machines directly to clarify data handling and deletion timelines before committing to production workflows.
Yes. Tinker supports fine-tuning on Inkling-Small for domain-specific tasks. Pricing for fine-tuning is not published; contact Thinking Machines for custom quotes. Fine-tuned models can be deployed on your infrastructure or served via Tinker API, enabling you to deliver specialized models to clients without exposing their training data to third parties.
For non-technical teams using Tinker Playground, rollout is same-day (sign up, start chatting). For engineers integrating Inkling-Small into production pipelines, expect 1-2 weeks for API integration, testing, and cost validation. Fine-tuning workflows add 2-4 weeks depending on dataset size and domain complexity. Self-hosting on your infrastructure adds 3-6 weeks for infrastructure setup and security review.