AI ToolAI Infrastructure

VLM Run

VLM Run Gateway consolidates access to 21+ visual models (OCR, detection, segmentation, video analysis, chat) behind a single OpenAI-compatible API, eliminating the need to manage separate integrations with OpenAI, Claude, Gemini, Qwen, and others.

VLM Run is an AI infrastructure platform, priced at $799/month on the Pro plan, integrating with OpenAI SDK, Zapier, MongoDB, and Claude Code. InnovaAI scores it 5.3/10 for agency resale.

Consider5.3/10

Agency Audit

VLM Run Gateway routes requests to 21+ visual models (OCR, detection, segmentation, video analysis) through a single OpenAI-compatible API, handling document chunking and video reassembly automatically. It outputs structured JSON schemas and integrates with OpenAI SDK, Zapier, Claude Code, and Pydantic AI. Best suited for agencies building document processing, video analysis, or physical AI workflows for clients. Resale potential exists for agencies serving construction, healthcare, or robotics verticals, but pricing is usage-based, making predictable client retainers harder to structure than fixed-tier SaaS.

ConsiderNo WLUsage Hybrid
Fit

5.3/10

Typical Margin

Depends on volume

Time-to-Value

3d about 3 days

Complexity
Low
Consider
Fit53
Visit VLM Run
Best For
  • You serve construction or healthcare clients who need document AI for blueprints, faxes, or clinical paperwork and can pass through token costs as usage charges.
  • Your agency builds internal AI tools and needs a unified API to avoid managing separate integrations with OpenAI, Claude, and Gemini for vision tasks.
  • You have robotics or physical AI clients requiring agentic data-labeling and can leverage MCP server deployment for agent-native vision workflows.
Not For
  • You want to resell a fixed-price retainer without tracking client token usage; VLM Run's per-token and per-operation pricing requires billing infrastructure or usage guardrails.
  • Your clients expect a white-labeled portal or branded interface; VLM Run does not offer a verified white-label program.
  • You need sub-second latency for real-time video processing; the platform is optimized for batch and asynchronous workflows, not live streaming.

Profit Path

Your Cost (USD)

$799/mo

Market Range

$3K–$8K/project

Revenue Model

Monthly Recurring

Planning benchmark at United States price levels. Not a measured market survey.

Platform Features

Core capabilities of VLM Run

21+ model routing via single endpoint

Route OCR, detection, segmentation, chat, and video analysis requests to 21+ visual models (OpenAI, Claude, Gemini, Qwen, Kimi, MiniMax, and others) through one OpenAI-compatible API. Agencies avoid maintaining separate integrations for each model provider.

Automatic document chunking and reassembly

Submit multi-page documents or long-form video in a single call; VLM Run chunks, processes, and reassembles outputs automatically. Reduces client-side orchestration work for document processing agencies.

Structured JSON schema enforcement

Enforce Pydantic schemas on every vision call, guaranteeing structured outputs for downstream workflows. Eliminates post-processing parsing logic and ensures consistent data formats across client projects.

MCP server for agent-native vision

Deploy VLM Run as an MCP server for Claude Code, Codex, and Pydantic AI agents. Agencies can build autonomous workflows that see, reason, and act on images and documents without custom API wrappers.

Video summarization and transcription

Summarize, transcribe, and search long-form video without manual review. Supports per-second billing for video generation and editing, enabling agencies to offer video analysis retainers to media and content clients.

Private VPC and on-premises deployment

Enterprise tier supports in-VPC deployments and on-premises installation. Agencies serving regulated industries (healthcare, finance) can offer compliant visual AI without data leaving client infrastructure.

What Makes VLM Run Different

Unique advantages vs similar tools in this niche

Cost-efficient document OCR at sub-cent per-page

vs Closed OCR APIs like Textract and Azure Doc AI

Pricing calculator shows gateway OCR costs $23-$90 per 100K pages vs $150-$1K for OCR APIs.

Handles 500-page PDFs or 2-hour videos in a single call

vs Building custom chunking and reassembly pipelines

Orchestration is built-in, eliminating pipeline development.

Open-weight models with no lock-in

vs Proprietary VLM APIs

Every model is open-weight and served through OpenAI-compatible API, allowing free model swapping.

Investment ROI Calculator

Value equation analysis for VLM Run, based on the Hormozi framework

What is the Hormozi framework? A four-factor score: (what the service delivers × how reliably it delivers) divided by (how long it takes × how much effort it requires). A higher Value Multiplier means a better return on the time and money invested: faster, easier, and more proven results.

Value MultiplierExcellent

2.8× value multiple: invest $799/mo and agencies typically charge $3K–$8K/project for the work it powers.

Outcome42
÷
Friction15

Why This Succeeds

Higher is better

Implementation Challenges

Lower is better

Strong ROI. VLM Run at $799/mo supports market rates of $3K–$8K. Its 2.8× value-equation score weighs client outcome and likelihood against the time and effort to deliver, not cost.

Best if:You serve construction or healthcare clients who need document AI for blueprints, faxes, or clinical paperwork and can pass through token costs as usage charges.Your agency builds internal AI tools and needs a unified API to avoid managing separate integrations with OpenAI, Claude, and Gemini for vision tasks.You have robotics or physical AI clients requiring agentic data-labeling and can leverage MCP server deployment for agent-native vision workflows.You need SOC 2, HIPAA, and BAA compliance for enterprise clients; the Pro plan ($799/mo) includes BAA and Zero-Data Retention, and Enterprise tier offers custom SLAs.

Pricing

VLM Run platform cost to your agency

Pro: $799/mo

Starter

Custom
  • Pay-as-you-go
  • Up to 10 requests/min
  • Community Discord
  • Basic Usage Logs

Pro

$799/mo
  • Up to 100 requests/min
  • Dedicated Slack support
  • Zero-Data Retention (ZDR)
  • Business Associate Agreement (BAA)
Enterprise

Enterprise

Custom
  • Invoiced billing
  • Tier-based pricing with volume discounts
  • Custom rate-limits
  • In-VPC deployments

How usage-based pricing works

VLM Run charges per consumption unit (per page grounding/confidence surcharge). Below are the component rates the vendor publishes. Each row is a separate charge: your total cost combines them based on your configuration and volume. Component rates range from $0.001 per page grounding/confidence surcharge.

Final agency cost = (sum of selected component rates) × client usage volume. Confirm a usage estimate with each client before quoting.

Component Rates

Cost per unit: total depends on your configuration and volume

Per page grounding/confidence surcharge
$0.001/ page grounding/confidence surcharge
Per code execution call
$0.001/ code execution call
Per image segment (fast/auto)
$0.01/ image segment (fast/auto)
Per page OCR/layout (fast/auto)
$0.01/ page OCR/layout (fast/auto)
Per image segment (pro)
$0.02/ image segment (pro)
Per image generate/edit (fast/auto)
$0.04/ image generate/edit (fast/auto)
Per page OCR/layout (pro)
$0.04/ page OCR/layout (pro)
Per second video generate/edit (fast/auto)
$0.15/ second video generate/edit (fast/auto)
Per image generate/edit (pro)
$0.24/ image generate/edit (pro)
Per 1M input tokens (vlmrun-orion-2:fast / auto)
$0.30/ 1M input tokens (vlmrun-orion-2:fast / auto)
Per second video generate/edit (pro)
$0.40/ second video generate/edit (pro)
Per 1M input tokens (kimi-2.6)
$0.66/ 1M input tokens (kimi-2.6)
Per 1M input tokens (gemini-flash-3.7)
$0.75/ 1M input tokens (gemini-flash-3.7)

Add-ons

Optional extras priced on top of any main plan

Add-on: 1M output tokens (vlmrun-orion-2:fast / auto)
$2.50
Add-on: 1M input tokens (vlmrun-orion-2:pro)
$1
Add-on: 1M output tokens (vlmrun-orion-2:pro)
$10
Add-on: 1M output tokens (kimi-2.6)
$3.41
Add-on: 1M input tokens (muse-spark-1.1)
$1.25
Add-on: 1M output tokens (muse-spark-1.1)
$4.25
Add-on: 1M output tokens (gemini-flash-3.7)
$3.75
Add-on: 1M input tokens (sonnet-5)
$2
Add-on: 1M output tokens (sonnet-5)
$10
Add-on: 1M input tokens (grok-4.5)
$2
Add-on: 1M output tokens (grok-4.5)
$6
Add-on: 1M input tokens (opus-4.8)
$5
Add-on: 1M output tokens (opus-4.8)
$25
Add-on: 1M input tokens (gpt-5.6-sol)
$5
Add-on: 1M output tokens (gpt-5.6-sol)
$30

No verified white-label program for VLM Run: client-facing delivery runs under the platform's native branding.

Reality Check

Trade-offs & Gotchas

Usage-based pricing (input/output tokens, per-page OCR, per-second video) makes client retainer pricing unpredictable. Agencies must either absorb cost variance or implement strict usage caps per client, adding operational complexity. No verified white-label program means client-facing surfaces display VLM Run branding.

Implementation Reality

Moderate effort: standard configuration with some customization needed

Effort: 3/10Time: 5/10

How This Accelerates White-Label Services

Who It's For

  • ai-infrastructure-providers
  • document-processing-agencies
  • video-analysis-agencies
  • enterprise-ai-teams

Acceleration Steps

  1. 1Create your account and complete setup wizard
  2. 2Configure route requests to 21+ visual models through one openai-compatible api
  3. 3Connect OpenAI SDK
  4. 4Launch your first client project

Academy for VLM Run

Work through it in order: the course for this service first, then the modules behind it.

Core concepts

The mental model you need to price and scope the work.

  1. Multi-Model Margin ShieldConcept

    Agencies integrating AI into client solutions face a hidden margin killer: lock-in to a single model provider. When one vendor raises prices or shifts capabilities, project feasibility and retainer margins erode overnight. The Multi-Model Margin Shield framework treats provider diversity as a financial hedge, not just a technical preference. By routing requests through an orchestration layer that can switch between Anthropic's Claude, OpenAI's GPT, and Google's Vertex AI based on cost and latency, agencies protect delivery margins and negotiate from strength. This approach also guards against capability shifts, such as when a model's safety guardrails change mid-project. For example, a recent study found GPT-6 Astra blocks 99.99% of direct prompt injections but fails 8.5% of hidden ones, while Claude Opus 5 performs differently, underscoring why redundancy matters for client-facing agents.

  2. Orchestration Layer Lock-InConcept

    Agencies integrating frontier models like Anthropic's Claude or OpenAI's GPT-5.6 into client solutions face a hidden risk: direct API dependency. Pricing changes, capability shifts, or outages at a single provider can erode project margins overnight. The framework of Orchestration Layer Lock-In argues that agencies should treat the model provider as a commodity and invest in a multi-model orchestration layer that abstracts routing, fallbacks, and cost management. This layer, exemplified by gateways like Helicone or OpenRouter, lets agencies switch between Claude, GPT, or others without rewriting client code. For instance, when Meta's ad AI altered approved creative post-launch, agencies relying on a single platform had no recourse; an orchestration layer would have enabled rapid failover to a safer model. By decoupling delivery from any one vendor, agencies protect margins and maintain negotiating power.

  3. Inference Cost EscalatorConcept

    The Inference Cost Escalator describes how an agency's AI infrastructure spend climbs silently as client projects scale. Each additional user, document, or agent loop multiplies token consumption, while premium model tiers (like Anthropic's Claude or OpenAI's GPT-5.6) carry higher per-token prices. Without a cost governance layer, a retainer that looked profitable at pilot stage can slip into negative margin as usage grows. Agencies can counter this by implementing a gateway that routes simple queries to cheaper models (e.g., Gemini 3.8 Flash) and reserves frontier models for complex reasoning, plus caching and rate limiting to cut redundant calls. For example, a recent benchmark comparing intelligence versus cost across models gives operators concrete data to match model tier to task complexity, preventing over-spend on routine work.

11 modules selected for VLM Run

Frequently Asked Questions

Answers about pricing, setup, implementation

VLM Run offers 3 pricing tiers, at $799/mo (Pro). Agencies typically achieve 55% profit margins when reselling to clients.

Starter is free with pay-as-you-go pricing and up to 10 requests/min. Pro is $799/month with up to 100 requests/min, dedicated Slack support, Zero-Data Retention, and BAA. Enterprise pricing is custom and includes invoiced billing, volume discounts, custom rate-limits, in-VPC deployments, SOC 2, HIPAA, BAA, and custom SLAs. Usage add-ons include per-token charges (input/output) ranging from $0.3 to $30 per 1M tokens depending on model, plus per-operation fees: $0.01 per page OCR (fast/auto), $0.04 per page OCR (pro), $0.15 per second video generation (fast/auto), and $0.4 per second video generation (pro).

No verified white-label program: client-facing surfaces display the VLM Run brand. Agencies can integrate VLM Run's API into their own applications or workflows, but cannot present a fully branded portal or interface to end clients.

Yes. VLM Run provides OpenAI-compatible endpoints, so the OpenAI SDK works natively. Zapier integration is also supported. Additional integrations include MongoDB, Claude Code, Codex, Pydantic AI, and Mastra, enabling agencies to embed VLM Run into multi-tool workflows.

Setup time depends on deployment model. API key provisioning is immediate (minutes). For agencies building client-facing applications, integration time ranges from hours (simple document OCR) to days (multi-model orchestration with custom schemas). Enterprise deployments with in-VPC or on-premises options require coordination with the VLM Run team.

Construction (blueprint and spec analysis), healthcare (fax and clinical paperwork processing), robotics and physical AI (agentic data-labeling), and enterprise AI teams building internal vision workflows. Document processing agencies and video analysis agencies are also strong fits.