Operating ProcedureExecution layer

QA Scorecard Calibration (QA)

A checklist with 7 steps: Pull a stratified sample of 40 to 60 scored interactions per client before trusting any automated score.

By InnovaAI ResearchPublished

What are the steps?

checklist

QA Scorecard Calibration (QA)

  1. 01

    Pull a stratified sample of 40 to 60 scored interactions per client before trusting any automated score

    Stratify by agent, channel, and call outcome so the sample reflects the full interaction mix rather than the easy-to-score calls. ScorebuddyCX auto-scores 100% of voice, chat, and email interactions, which makes a manual calibration sample the only way to verify the model is grading the way the client's program intends.

  2. 02

    Have two reviewers independently score the same 20 interactions against the client's rubric

    Measure inter-rater agreement before comparing human scores to machine scores. If the two humans disagree on more than 15% of items, the rubric is ambiguous and no AI score can be trusted yet.

  3. 03

    Compare automated scores against the human consensus and flag every variance above 10 points

    Group the variances by scorecard category (greeting, compliance disclosure, resolution, upsell attempt) rather than by agent. Category clustering tells you whether the model misreads a specific behavior or is drifting across the board.

  4. 04

    Rewrite any rubric line that a reviewer cannot apply consistently in under 30 seconds

    Vague lines like 'demonstrated empathy' produce noisy scores in every platform. Replace them with observable behaviors: 'acknowledged the customer's stated problem in the agent's first two sentences.'

  5. 05

    Re-run the calibration sample after rubric edits and document the score delta

    A single re-run pass is enough to confirm the edits moved scores in the intended direction. Record the before-and-after numbers in the client's QA log so the change is auditable at renewal.

  6. 06

    Set a monthly recalibration cadence tied to the client's coaching cycle

    Score drift accumulates as call types, promotions, and scripts change. Tying recalibration to the coaching cycle keeps the scorecard aligned with what agents are actually being coached on that month.

  7. 07

    Publish the calibration summary to the client with the variance data attached

    Clients who see the variance numbers alongside the scores treat QA as a measurement program rather than a black box. This is also the artifact that protects the agency when an agent disputes a score.