Operating ProcedureExecution layer

Confidence Threshold Calibration (QA)

A checklist with 7 steps: Pull 200 closed tickets from the client's last 90 days and label each by true intent and true urgency.

By InnovaAI ResearchPublished

What are the steps?

checklist

Confidence Threshold Calibration (QA)

  1. 01

    Pull 200 closed tickets from the client's last 90 days and label each by true intent and true urgency

    Use the client's own resolution notes as ground truth rather than agent memory. A 200-ticket sample is large enough to expose rare intents without turning the exercise into a two-week project.

  2. 02

    Score the classifier's top-1 prediction against that labeled set and record precision and recall per intent bucket

    Report per-bucket numbers, not a single blended accuracy figure. A model that hits 91% overall can still miss billing disputes at 60%, and billing disputes are the tickets clients call about.

  3. 03

    Set the auto-route threshold where precision clears 95% for that bucket, and route everything below it to a human queue

    Thresholds belong per intent, not globally. Julia 1's design, which selects among 2 to 20 supplied answer options across 52 locales, makes per-locale thresholds practical where a single global cutoff would not be.

  4. 04

    Define a separate escalation threshold for sentiment and urgency signals that overrides the routing threshold

    A refund request with a 99% intent match still needs a human if the customer has written three times in an hour. Sentiment and repeat-contact count should be able to force escalation regardless of classification confidence.

  5. 05

    Write the fallback path for low-confidence tickets and confirm it lands in a staffed queue during the client's business hours

    An unmonitored fallback queue is worse than no automation, because the ticket looks handled in the dashboard while the customer waits. Verify queue ownership with the client's support lead in writing.

  6. 06

    Run a two-week shadow period where the model predicts but humans still route, then compare predicted versus actual routing

    Shadow mode surfaces threshold errors before they reach customers. Track the disagreement rate weekly and adjust thresholds once, not continuously, so the comparison stays clean.

  7. 07

    Document the calibration date, sample size, and threshold values in the client's retainer file and schedule the next review

    Intent mix shifts with product launches and seasonal demand, so thresholds decay. A quarterly recalibration cadence keeps the numbers defensible when the client asks why deflection moved.