Operating ProcedureExecution layer

Production Trace Triage (QA)

A sequence with 7 steps: Define the failure taxonomy before opening any trace viewer.

By InnovaAI ResearchPublished

What are the steps?

sequence

Production Trace Triage (QA)

  1. 01

    Define the failure taxonomy before opening any trace viewer

    Agree on 5 to 7 named failure classes with the client: wrong answer, hallucinated citation, tool call error, latency breach, cost overrun, tone drift, and refusal. Every triage decision maps back to one of these labels.

  2. 02

    Pull the last 7 days of traces filtered to production traffic only

    Exclude staging and internal test runs. A platform like Langfuse or Arize can segment by environment tag so the sample reflects what real users experienced, not what your team rehearsed.

  3. 03

    Sort by cost and latency percentiles before reading any output text

    The p95 and p99 buckets surface the 3 to 5 percent of sessions that generate most client complaints and most token spend. Read those first; the median is usually fine and rarely the source of churn.

  4. 04

    Score a stratified sample of 50 traces against the taxonomy

    Take 50 sessions weighted toward high-cost and high-latency buckets. Braintrust and Confident AI both support LLM-as-judge scoring, but a human reviewer should label at least 10 of the 50 to calibrate the judge.

  5. 05

    Trace each flagged failure back to its root span

    A wrong answer often originates in a retrieval step, not the generation step. Follow the span tree until you find the first point where the output diverged from expected behavior.

  6. 06

    Log every confirmed failure as a regression test case

    Convert the trace into a repeatable eval input with the expected output attached. This turns a one-off incident into a permanent guardrail that runs on every future deployment.

  7. 07

    Deliver a triage memo to the client within 48 hours

    Include failure counts by class, the two highest-impact root causes, and a remediation estimate in hours. Clients tolerate known failures far better than unexplained ones, and the memo becomes the basis for a change order.