Failure PatternDecision layer
The Weave Router Quota Drain Trap: Why Agencies Fail With Weave Router on Fixed-Fee Retainers
Symptom: Client invoices show frontier-model token charges even though the router was sold as a cost-control layer, because flat-rate subscription quota was never drained before per-token billing kicked in. Root cause: Agencies configure routing rules against a generic workload pattern instead of the client's real coding mix, so the complexity classifier has no calibrated threshold and defaults to escalation.
By InnovaAI ResearchPublished
How do you recognize it?
- •Client invoices show frontier-model token charges even though the router was sold as a cost-control layer, because flat-rate subscription quota was never drained before per-token billing kicked in
- •Routine turns (renames, formatting, small edits) land on Claude or GPT while DeepSeek, Llama, and Gemini sit idle, so the routing policy is effectively bypassed
- •Agents loop on a failing test or stalled refactor and keep escalating to frontier models turn after turn instead of being caught by the escalation classifier
- •The $1,800 Weave Router Starter Setup was delivered in 16h, but no usage dashboard report template was wired to the client's actual workload, so nobody can show where spend went
- •Delivery leads discover the client's Claude Code or Cursor harness was never detected by the npx @workweave/router command, meaning turns were never routed at all
Why does it happen?
- •Agencies configure routing rules against a generic workload pattern instead of the client's real coding mix, so the complexity classifier has no calibrated threshold and defaults to escalation
- •Cache-aware cost checks and subscription quota draining are left disabled at go-live, which means the router pays per-token for work the client's flat-rate plan already covers
- •The escalation classifier is treated as a safety net rather than a tuned control, so stalled or looping tasks are allowed to burn frontier turns before anyone intervenes
- •Weave Router is infrastructure, not a client-facing deliverable, and agencies sell it as a standalone service without owning the ongoing routing policy, leaving no retainer scope for tuning after handoff
How do you fix it?
- •Open the router config and verify the npx @workweave/router detection pass actually registered Claude Code, Codex, and Cursor before enabling live traffic
- •Enable cache-aware cost checks and subscription quota draining, then re-run a sample of the client's last 50 agent turns to confirm cheap models are being selected first
- •Re-tune the routing policy thresholds against the client's real workload: send routine turns to DeepSeek, Llama, or Gemini and reserve Claude and GPT for tasks the escalation classifier flags
- •Rebuild the usage dashboard report template around model spend and routing decisions per turn, and attach it to the retainer as a monthly deliverable so tuning has a billable owner
More on Weave Router
- StrategyWhy Weave Router Is a Cost-Control Layer, Not an Agency Retainer
- ConceptWeave Router Escalation Budget
- Evaluation RuleWeave Router Rule: Adopt Only When Your Client's Coding-Agent Token Spend Exceeds the $1,800 Setup Fee
- Decision FrameworkWeave Router: Buy vs Skip (Agency AI Coding Spend)
- Implementation BlueprintWeave Router Cost-Control Retainer Build (5-7 days)
- Operating ProcedureWeave Router Client Routing Policy Setup (Onboarding)
More for AI Code Tools
- Failure PatternsThe Unreviewed Merge Trap: Why AI Code Tools Fail in Agency Delivery
- Failure PatternsThe Scaffolding-Only Trap: Why AI Code Tools Stall in Agency Delivery
- Failure PatternsThe Leadcode Credential Drift Trap: Why Agencies Fail With Leadcode in Multi-Client Delivery
- Failure PatternsWhy Agencies Fail With GenerativeIDE in Security-Sensitive Client Work