Failure PatternDecision layer

The Weave Router Quota Drain Trap: Why Agencies Fail With Weave Router on Fixed-Fee Retainers

Symptom: Client invoices show frontier-model token charges even though the router was sold as a cost-control layer, because flat-rate subscription quota was never drained before per-token billing kicked in. Root cause: Agencies configure routing rules against a generic workload pattern instead of the client's real coding mix, so the complexity classifier has no calibrated threshold and defaults to escalation.

By InnovaAI ResearchPublished

How do you recognize it?
  • •Client invoices show frontier-model token charges even though the router was sold as a cost-control layer, because flat-rate subscription quota was never drained before per-token billing kicked in
  • •Routine turns (renames, formatting, small edits) land on Claude or GPT while DeepSeek, Llama, and Gemini sit idle, so the routing policy is effectively bypassed
  • •Agents loop on a failing test or stalled refactor and keep escalating to frontier models turn after turn instead of being caught by the escalation classifier
  • •The $1,800 Weave Router Starter Setup was delivered in 16h, but no usage dashboard report template was wired to the client's actual workload, so nobody can show where spend went
  • •Delivery leads discover the client's Claude Code or Cursor harness was never detected by the npx @workweave/router command, meaning turns were never routed at all
Why does it happen?
  • •Agencies configure routing rules against a generic workload pattern instead of the client's real coding mix, so the complexity classifier has no calibrated threshold and defaults to escalation
  • •Cache-aware cost checks and subscription quota draining are left disabled at go-live, which means the router pays per-token for work the client's flat-rate plan already covers
  • •The escalation classifier is treated as a safety net rather than a tuned control, so stalled or looping tasks are allowed to burn frontier turns before anyone intervenes
  • •Weave Router is infrastructure, not a client-facing deliverable, and agencies sell it as a standalone service without owning the ongoing routing policy, leaving no retainer scope for tuning after handoff
How do you fix it?
  • •Open the router config and verify the npx @workweave/router detection pass actually registered Claude Code, Codex, and Cursor before enabling live traffic
  • •Enable cache-aware cost checks and subscription quota draining, then re-run a sample of the client's last 50 agent turns to confirm cheap models are being selected first
  • •Re-tune the routing policy thresholds against the client's real workload: send routine turns to DeepSeek, Llama, or Gemini and reserve Claude and GPT for tasks the escalation classifier flags
  • •Rebuild the usage dashboard report template around model spend and routing decisions per turn, and attach it to the retainer as a monthly deliverable so tuning has a billable owner