Production Trace Triage (QA)
A sequence with 7 steps: Define the failure taxonomy before opening any trace viewer.
By InnovaAI ResearchPublished
What are the steps?
Production Trace Triage (QA)
- 01
Define the failure taxonomy before opening any trace viewer
Agree on 5 to 7 named failure classes with the client: wrong answer, hallucinated citation, tool call error, latency breach, cost overrun, tone drift, and refusal. Every triage decision maps back to one of these labels.
- 02
Pull the last 7 days of traces filtered to production traffic only
Exclude staging and internal test runs. A platform like Langfuse or Arize can segment by environment tag so the sample reflects what real users experienced, not what your team rehearsed.
- 03
Sort by cost and latency percentiles before reading any output text
The p95 and p99 buckets surface the 3 to 5 percent of sessions that generate most client complaints and most token spend. Read those first; the median is usually fine and rarely the source of churn.
- 04
Score a stratified sample of 50 traces against the taxonomy
Take 50 sessions weighted toward high-cost and high-latency buckets. Braintrust and Confident AI both support LLM-as-judge scoring, but a human reviewer should label at least 10 of the 50 to calibrate the judge.
- 05
Trace each flagged failure back to its root span
A wrong answer often originates in a retrieval step, not the generation step. Follow the span tree until you find the first point where the output diverged from expected behavior.
- 06
Log every confirmed failure as a regression test case
Convert the trace into a repeatable eval input with the expected output attached. This turns a one-off incident into a permanent guardrail that runs on every future deployment.
- 07
Deliver a triage memo to the client within 48 hours
Include failure counts by class, the two highest-impact root causes, and a remediation estimate in hours. Clients tolerate known failures far better than unexplained ones, and the memo becomes the basis for a change order.