Pre-Launch Eval Gate (Onboarding)
A checklist with 7 steps: Freeze the client's acceptance criteria in writing before any prompt work starts.
By InnovaAI ResearchPublished
What are the steps?
Pre-Launch Eval Gate (Onboarding)
- 01
Freeze the client's acceptance criteria in writing before any prompt work starts
Convert vague asks like 'sounds on brand' into 8 to 12 scored criteria with pass thresholds, then get the client to sign off. This becomes the contract you test against at delivery.
- 02
Build a golden dataset of 30 to 50 real client inputs, not synthetic ones
Pull actual support tickets, sales emails, or call transcripts. Synthetic prompts hide the messy edge cases that cause production failures.
- 03
Stand up tracing on day one so every call is logged from the first test run
Platforms such as Langfuse and Arize capture hierarchical traces of each LLM call, tool invocation, and retrieval step, which gives you a replayable record when something breaks.
- 04
Score baseline outputs with an LLM-as-judge rubric before tuning anything
Braintrust and Confident AI both support judge-based scoring against custom criteria. Record the baseline number; without it you cannot prove improvement to the client.
- 05
Run adversarial and red-team cases against the agent before it touches client data
Confident AI bundles red teaming against adversarial attacks. Recent incidents where agents breached external systems make this a client-trust requirement, not a nice-to-have.
- 06
Set cost and latency ceilings per interaction and alert on breach
Define a dollar cap per conversation and a p95 latency target. Hidden cost spikes are one of the failure modes that erodes retainer margins fastest.
- 07
Document the eval harness so a second delivery team can rerun it without you
Version the dataset, rubric, and thresholds in the client folder. This is what makes the engagement portable and defensible at handoff.