Operating ProcedureExecution layer

Model Routing and Fallback Gate (Delivery)

A sequence with 8 steps: Inventory every model call the client deliverable depends on before writing routing logic.

By InnovaAI ResearchPublished

What are the steps?

sequence

Model Routing and Fallback Gate (Delivery)

  1. 01

    Inventory every model call the client deliverable depends on before writing routing logic

    List each call site, the model currently attached to it, and the business consequence if that model is deprecated, repriced, or degraded. Anthropic's Claude Haiku 5.5 shipped with a 1 million token context window at $0.10 per million input tokens, so a call site priced against an older small model may already be overpaying by an order of magnitude.

  2. 02

    Classify each call site as latency-bound, cost-bound, or quality-bound

    A chat widget answering pricing questions is latency-bound; a nightly batch that summarizes 4,000 support tickets is cost-bound; a contract redline assistant is quality-bound. The classification decides which fallback order is acceptable when the primary provider is unavailable.

  3. 03

    Route all traffic through one gateway layer rather than calling providers directly from application code

    Gateways such as Helicone, Portkey, and TrueFoundry proxy requests so caching, rate limits, retries, and per-request logging live in one place. When a client asks why last Tuesday's invoice spiked, the answer comes from a dashboard instead of a grep through application logs.

  4. 04

    Define a primary and at least two fallback models per call site, with a written trigger for switching

    Triggers should be mechanical: error rate above 2% over five minutes, p95 latency above the client's stated ceiling, or a published price change. OpenRouter and similar routing services can hold the fallback list, but the switch conditions must be documented in the delivery file, not left to whoever is on call.

  5. 05

    Run a paired evaluation on the fallback model before it goes live for any client

    Send 50 to 100 real (redacted) prompts through both models and score outputs against the acceptance criteria in the statement of work. A September 2026 technical analysis found that top benchmark scores do not reliably predict production performance, so a leaderboard rank is not a substitute for a paired run on the client's own inputs.

  6. 06

    Set a spend ceiling and an alert threshold per client per week

    CostPerPrompt and comparable trackers exist because token spend drifts quietly. A retainer priced at $6,000 per month cannot absorb an unattended $2,400 inference bill, and the alert needs to fire before the invoice, not after.

  7. 07

    Log the routing decision, model version, and token count for every production call

    Retain the log for the length of the client contract. When a client disputes an output six weeks later, the log is the only way to reconstruct which model produced it and under what prompt version.

  8. 08

    Review the routing map at each monthly delivery check-in and reprice any call site whose model changed

    Model releases land weekly. OpenAI's GPT-6 family alone spans three variants tuned for prototyping, feature development, and multi-step orchestration, and attaching the wrong tier to a workflow raises cost while lowering output quality. The review keeps the map current without a rebuild.