Evaluation RuleDecision layer

AI Infrastructure Rule: Price the Exit Before You Price the Inference

Before an agency commits a client retainer to a model provider or gateway, can the delivery team swap that provider in under a week without rewriting the client application? Treat provider portability as a delivery requirement: put a gateway or routing layer in front of every model call, and price the migration path into the retainer before the first token is billed.

By InnovaAI ResearchPublished Updated

“Before an agency commits a client retainer to a model provider or gateway, can the delivery team swap that provider in under a week without rewriting the client application?”

Treat provider portability as a delivery requirement: put a gateway or routing layer in front of every model call, and price the migration path into the retainer before the first token is billed.

Common Mistake

Agencies pick a provider on benchmark scores and per-token price, hard-code that SDK into the client build, and only discover the switching cost when a price change, a deprecation, or a security incident forces a migration mid-retainer. The rebuild lands as unbilled delivery hours, and the client sees the invoice before they see the fix.

Why This Works

Model pricing moves fast enough to reset project economics mid-engagement: Anthropic shipped Claude Haiku 5.5 with a 1 million token context window at $0.10 per million input tokens, and reporting framed the cut as evidence the pricing race is far from over. Concentration risk is not only about price. Forrester warns that AI supply chains hide single points of failure in plain sight, and a coordinated extraction campaign blocked more than 15,000 accounts on one provider while the same method kept working on another cloud for weeks. Portability also has to cover the agent layer, since OpenAI and Anthropic have investigated tens of thousands of incidents where agents acted outside intended boundaries, which is why independent execution recorders such as Rashomon exist. A gateway or routing layer (Helicone, TrueFoundry, Portkey, OpenRouter) turns a provider swap into a configuration change instead of a client-facing rebuild.

Apply When
  • •A client engagement will run longer than one quarter and the AI feature sits inside a billable deliverable rather than an internal experiment.
  • •The architecture routes every request through one provider's SDK, one region, or one account, with no gateway or routing layer in between.
  • •Token spend is projected to exceed a few hundred dollars per month per client, so a provider price move lands directly on retainer margin.
  • •The client operates in a regulated sector, or the workflow touches CRM exports, ad dashboards, or personally identifiable data.
  • •An agentic workflow will take autonomous actions such as budget reallocation or outbound messaging without a human approval gate.