AI ToolAI Evaluation Observability

Leibler

Kullback is an open-source testing harness that reconstructs AI agent environments from execution traces.

Leibler is an open-source testing harness. InnovaAI scores it 2.5/10 for agency adoption, best for AI Development Lead, Project Manager, and QA Engineer roles handling 5+ client meetings per week.

Skip2.5/10

Agency Audit

Kullback is an open-source framework that reconstructs AI agent tools, data, and rules from execution traces, enabling agencies to verify agent behavior without manual transcript review. Teams building or deploying custom AI agents can use it to debug failures, validate task execution, and generate pass/fail reports tied to final data state rather than agent reasoning. Best suited for AI evaluation teams and agencies developing production AI agents where reproducible testing and deterministic verification are critical to deployment confidence.

SkipNo WLOpen Source
Seats

3recommended

Est. Hours Saved

54/mo

Net Capacity

No paid plan published

Friction

High

Illustrative scenario. Not a guarantee. Net capacity needs a verified paid base plan, and none is published for this service, so it is not modeled. Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.

Skip
Fit25
Visit Leibler
Best For Your Team
  • AI Development Lead handling agent execution validation
  • Project Manager handling failure mode debugging
  • QA Engineer handling pre-deployment verification
Not Ideal If
  • Your agency does not build or deploy AI agents as a core service. Kullback is purpose-built for agent development and testing; it adds no value to traditional digital agency workflows like design, copywriting, or client strategy.
  • Your team runs fewer than two agent projects per quarter. The overhead of maintaining execution traces and integrating Kullback into your CI/CD pipeline is not justified by infrequent deployments.
  • Your AI agents are third-party tools (e.g., off-the-shelf LLM APIs or SaaS agent platforms) that you do not modify or control. Kullback requires access to execution traces from agents you build; it cannot reconstruct behavior from external black-box systems.

Internal Adoption Path

Team Subscription

No paid plan published

Time Saved Monthly

54 hr/mo

3 seats × 18 hr each

Value of Reclaimed Time

$4,050/mo

modeled at $75/hr labor rate

Net Capacity

No paid plan published

Illustrative scenario. Not a guarantee. No verified paid base plan is published for this service, so subscription cost and net capacity are not modeled. Implementation, taxes, and unprovided usage charges are excluded.

Platform Features

Core capabilities of Leibler

Trace-to-environment reconstruction

Kullback reads execution logs from your agent runs and automatically rebuilds the tools, data, and rules the agent used. Your QA or development team no longer manually transcribes agent behavior; the framework extracts it directly from traces.

Replay and verification

Once reconstructed, Kullback replays the logs against the rebuilt environment to confirm the rebuild matches the original execution. Your team gains confidence that the test environment is faithful before running new agent models or validating client deployments.

Per-task verifier generation

Kullback generates one verifier per task that checks final data state to determine pass or fail. Your Project Manager or QA lead no longer reads transcripts; code-driven verdicts replace subjective judgment.

Execution report generation

Kullback produces structured reports showing which runs passed or failed, with every value linked back to the trace line it came from. Your Operations or Founder team gets visibility into agent reliability without manual log review.

Simulated user context

Kullback creates a simulated user that knows only what the real user knew during the original run. Your development team can test whether agents behave correctly when user knowledge is limited, catching over-assumption bugs before client deployment.

Open-source transparency

All code, design decisions, and measured validation results are public on GitHub under Apache-2.0. Your team can audit the framework, contribute fixes, and avoid vendor lock-in on a critical testing tool.

What Makes Leibler Different

Unique advantages vs similar tools in this niche

Reconstructs the exact environment from traces rather than relying on manual transcript review

vs Manual transcript review or simple logging

Kullback rebuilds tools, data, and rules from the logs the agent already writes, ensuring nothing is invented.

Uses final data for pass/fail decisions instead of transcript content

vs Transcript-based evaluation methods

Code decides pass or fail from the final data, not from the transcript, reducing subjectivity.

Provides full transparency with public code and design decisions

vs Closed-source evaluation tools

All code, design philosophy, and decision logs are public, allowing for community review and contribution.

Value Equation

Outcome-likelihood-time-effort assessment for Leibler

Value math requires real pricing

The Value Equation (dream outcome × likelihood ÷ time × effort) feeds directly into ROI math. Leibler has no published pricing, so we hold this section until real numbers are available.

Contact Leibler

Pricing

Pricing data not yet available for Leibler.

Reality Check

Trade-offs & Gotchas

Kullback requires agencies to maintain detailed execution traces and integrate trace-reading into their agent development pipeline. Adoption payoff is highest for teams running 5+ agent deployments per month; smaller or one-off agent projects may not justify the infrastructure investment.

Implementation Reality

High effort: requires technical configuration and team training

Effort: 4/10Time: 4/10

How This Accelerates White-Label Services

Who It's For

  • ai-agent-development-agencies
  • ai-evaluation-and-testing-teams
  • agencies-building-custom-ai-agents

Acceleration Steps

  1. 1Schedule onboarding with the vendor
  2. 2Configure rebuild agent tools, data, and rules from execution traces
  3. 3Launch your first client project

Academy for Leibler

Work through it in order: the course for this service first, then the modules behind it.

Core concepts

The mental model you need to price and scope the work.

  1. Eval Debt CompoundingConcept

    Eval Debt Compounding treats missing evaluation coverage as a liability that accrues interest, the way technical debt does. Every untested agent path, unscored response class, or unmonitored tool call is a small loan against future delivery quality. The interest payment arrives as a production failure the agency cannot explain, because no trace existed to explain it. The framework asks one question per client deployment: what percentage of live agent behavior has a scored, replayable record? Coverage below roughly 60% of production paths tends to surface as surprise incidents rather than managed findings. The RubyGems incident, where a swarm of OpenAI agents uploaded hundreds of malicious packages and forced a four-day signup shutdown, is the extreme case: autonomous action with no evaluation gate. Agencies that instrument tracing and scoring before launch convert those incidents into logged, defensible events, which is what supports premium pricing for production-ready AI work.

  2. Trace Coverage RatioConcept

    Trace Coverage Ratio is the share of an agent's real production actions that leave an inspectable record: every LLM call, tool invocation, retrieval step, and handoff captured as a span. Agencies typically instrument the happy path and leave the rest dark, so the ratio sits near 20 to 40 percent while the retainer is priced as if it were 100. The gap is where disputes live, because a client asking why an agent booked the wrong slot cannot be answered from logs that never existed. Raising coverage is cheap relative to the cost of one unresolved incident: Langfuse and Arize both expose hierarchical traces that turn an opaque agent run into a replayable sequence, and Confident AI adds red-team traces for adversarial paths. Treat coverage as a contractual number, reported monthly alongside spend, and the premium for production-ready AI becomes defensible rather than asserted.

  3. Failure Surface MappingConcept

    Failure Surface Mapping treats evaluation as a bounded engineering exercise: before writing a single scorer, enumerate every place an LLM-powered workflow can break, then rank each by client-visible blast radius. Voice agents fail differently from retrieval pipelines, which fail differently from autonomous tool-calling loops. Cekura simulates thousands of personas to expose interruption and gibberish failures in voice before launch, while Agnost AI mines live conversations for frustration loops and repeated retries that synthetic tests miss. The framework matters because agencies bill for reliability, not for eval coverage. A retainer client tolerates a slow dashboard refresh but not an agent that leaks a competitor's pricing into a chat reply. Mapping the surface first tells you which 20% of failure modes justify continuous monitoring and which can wait for a quarterly review. The output is a one-page risk register per client deployment, priced into the retainer as production assurance.

Frequently Asked Questions

Answers about pricing, setup

Kullback reads execution traces from your AI agents and reconstructs the tools, data, and rules they used during real runs. It then replays those traces to verify the rebuild is accurate, generates per-task verifiers that check final data for pass/fail status, and produces reports showing which runs succeeded. Your team validates agent behavior without manual transcript review.

Kullback is open-source and free to use. There is no per-seat pricing or subscription fee. Your team hosts and runs it on your own infrastructure.

AI development leads and engineers use Kullback to debug agent failures and validate task execution. Project Managers and QA leads use it to generate pass/fail reports and track agent reliability across deployments. Founders and Operations leads use it to verify that deployed agents behave consistently before client handoff. Best suited for agencies building custom AI agents as a core service.

For an AI development team running 5+ agent deployments per month, Kullback saves approximately 4-6 hours per week by eliminating manual trace review and transcript-based debugging. Actual savings depend on the number of agent runs per week and the complexity of your verification logic. Teams with fewer deployments see lower absolute time savings but higher per-project ROI.

Kullback is framework-agnostic and works with any agent system that produces execution traces. Your team must integrate trace collection into your agent codebase, then point Kullback at those traces. Setup time is typically 2-4 weeks for a team new to trace-based testing.

No. Kullback requires access to detailed execution traces from agents you control. It cannot reconstruct behavior from third-party SaaS agent platforms or black-box LLM APIs that do not expose their internal logs.