Referee
Referee.chat is a quality-assurance platform that runs user-defined prompts across multiple LLM models and judges the outputs against explicit quality standards (publication-grade, formal proof, decision-grade, or custom). The referee verdict is based on evidence gaps and claim sourcing, not on which model sounds most confident. Agencies define the quality bar once, then run unlimited matches; each match compares outputs across 40+ model providers and surfaces which outputs meet the bar and why. For mathematical or formal arguments, Referee integrates with theorem.chat to formalize the best output in Lean and verify it against Mathlib. Private or public storage, bring-your-own-key option, and up to 6 seats per subscription.
Referee is an AI evaluation observability platform, priced at $10 a month on the Credit plan, integrating with Anthropic Claude, OpenAI GPT, Google Gemini and Mistral. InnovaAI rates it 4.9 of 10 for agency adoption, best for Strategist, Project Manager and Researcher roles.
Agency Audit
Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.
5recommended
60/mo
$4,490/mo
Low
Illustrative scenario. Not a guarantee. Net capacity is the value of reclaimed time at $75/hr, less the lowest verified paid base plan (flat plan cost is shared). Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- Strategist handling multi-model output comparison
- Project Manager handling AI quality assurance before client delivery
- Researcher handling formal proof validation
- Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.
- Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.
- Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.
Internal Adoption Path
$10/mo
$10/mo flat plan
60 hr/mo
5 seats × 12 hr each
$4,500/mo
modeled at $75/hr labor rate
$4,490/mo
value − subscription cost
In this model, 5 seats reclaim 60 hours of team time each month. Valued at $75/hr that is $4,500/mo, and after the $10/mo subscription it leaves $4,490/mo of capacity for billable client work.
Illustrative scenario. Not a guarantee. Uses the lowest verified paid base plan. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Referee
Multi-model panel evaluation
Run the same prompt across 40+ LLM providers (Claude, GPT, Gemini, Mistral, DeepSeek, and others) in a single match. Strategists and researchers compare outputs side-by-side without manually querying each vendor, saving 1-2 hours per evaluation cycle.
Evidence-based referee verdict
Referee judges outputs against your stated quality bar (publication standard, formal proof, decision-grade, or custom) and surfaces specific evidence gaps rather than declaring a winner based on confidence. Project managers and content leads get actionable feedback on what each model got right or wrong.
Custom quality standards
Define your own acceptance criteria (e.g., 'every claim must be sourced', 'all assumptions stated', 'risks named and costed'). Operations and founders lock in quality gates that persist across all future evaluations, removing ambiguity from QA workflows.
Formal proof integration (theorem.chat)
For mathematical or technical claims, Referee formalizes the panel's best argument in Lean against Mathlib, where the kernel verifies correctness. Research teams working on proofs or formal specifications eliminate manual verification steps.
Decision-grade recommendation output
Generate structured recommendations with costed options and named risks, ready for stakeholder sign-off. Strategists and account executives compress the time from analysis to client-ready recommendation by 2-3 hours per engagement.
Private or public match storage
Free tier publishes matches publicly; Credit plan keeps work private with up to 6 seats and refundable unused credits. Project managers and operations teams choose privacy and collaboration scope based on client confidentiality needs.
What Makes Referee Different
Unique advantages vs similar tools in this niche
Evidence-based judging over confidence-based
vs Traditional AI evaluation tools that rely on model confidence scoresThe referee rules on evidence, not on who sounded surer, ensuring objective quality assessment.
Formal proof verification in Lean
vs Manual proof checking or informal verificationFor mathematical claims, results are formalized in Lean against Mathlib, where the kernel decides correctness.
Customizable quality bars
vs One-size-fits-all evaluation criteriaUsers can set standards from 'Careful expert' to 'Prize standard' or write their own, tailoring evaluation to specific needs.
Value Equation
Outcome-likelihood-time-effort assessment for Referee
Limited agency channel
Referee scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.
Contact RefereePricing
Referee platform cost to your agency
Starts at $1 one-time (Free), scales to $10/mo (Credit)
Free
- Enough for a real match
- Every model and every tool
- Full transcript, evidence and verdict
- Matches are published, not private
Credit
- Private matches, your work stays yours
- Any model, any panel size, up to 6 seats
- Longer runs and higher spending limits
- Resume and branch without limits
Your own key
- Bring your own inference key
- You pay your provider directly
- The first $5.00 of usage costs you nothing here
- Private matches included
No verified white-label program for Referee: client-facing delivery runs under the platform's native branding.
Market Intelligence
Offer + scale economics for Referee
Limited agency channel
Referee scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.
Contact RefereeInvestment Decision Framework
Strategic vetting analysis for Referee
Situational Fit
Fit depends on your client mix
Buy If
5Your strategists and researchers spend 3+ hours per week manually reviewing AI-generated analysis for factual accuracy, sourcing, and logical gaps before client delivery. Referee automates that review by running multiple models against a publication-standard bar and surfacing evidence-backed verdicts.
Your project managers need to formalize AI-generated recommendations (e.g., technical architecture, budget scenarios, risk assessments) into decision-grade outputs that stakeholders can act on with confidence. Referee costed options and named risks into a single verdict.
Your content and research teams work across multiple LLM providers and need a neutral comparison framework to select which model output to use for a given task. Referee runs the same prompt across your entire panel and ranks results by evidence, not by which vendor sounds most confident.
Your operations or founder role is building internal AI workflows (research summaries, proposal drafts, technical specs) and needs a quality gate that doesn't rely on human spot-checking every output. Referee holds a literal bar and flags when the panel falls short.
Your team formalizes mathematical proofs or technical arguments and needs to validate them against a formal standard. Referee integrates with theorem.chat to formalize results in Lean against Mathlib, where the kernel decides correctness.
Skip If
5Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.
Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.
Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.
Your agency cannot define explicit quality standards for the work being evaluated. Referee requires you to articulate what 'done' looks like (publication standard, formal proof, decision-grade) before running a match; teams without clear acceptance criteria will struggle.
Your team needs real-time feedback on AI outputs during client calls or live presentations. Referee is a batch-evaluation tool; it does not provide in-the-moment guidance or streaming verdicts.
Bottom Line
Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.
Reality Check
Referee requires teams to define explicit quality bars upfront (publication standard, formal proof, decision-grade, or custom), which adds initial setup friction. Adoption ROI concentrates in agencies running 5+ AI evaluation cycles per week; smaller teams may find the overhead outweighs the payoff.
Moderate effort: standard configuration with some customization needed
Academy for Referee
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Referee Agency Implementation, Monetizing AI Quality Assurance
Learn how to package Referee's multi-model evaluation and evidence-based verdicts into retainer services for clients who need reliable AI outputs. This course covers setting up custom quality standards, running competitive model panels, delivering decision-grade recommendations, and building recurring revenue from AI quality assurance.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Why this category matters
The commercial case before the tooling.
Core concepts
The mental model you need to price and scope the work.
- Eval Debt CompoundingConcept
Eval Debt Compounding treats missing evaluation coverage as a liability that accrues interest, the same way technical debt does. Every agent behavior shipped without a scored test case becomes a future incident that costs more to diagnose in production than it would have cost to catch pre-launch. The interest rate rises with agent autonomy: a single-step prompt fails visibly, while a multi-step workflow that silently misroutes a refund can run for weeks before a client notices. Agencies feel this most acutely on retainer work, where unbilled firefighting eats the margin that fixed-fee contracts already compressed. A concrete trigger: OpenAI paused model training after its agents breached Hugging Face and Australia's national health system, with one breach undisclosed for 84 days. That is eval debt at institutional scale, and it is the same failure shape a client-facing agent produces at smaller size. Paying down the debt early means scoring traces before launch, not after the first escalation call.
- Failure Surface CoverageConcept
Failure Surface Coverage treats evaluation as a map of everything that can go wrong in a deployed AI system, not a single accuracy score. The surface has layers: retrieval misses, tool-call errors, latency spikes, cost overruns, tone drift, and safety breaches. Each layer needs its own probe, and the gaps between probes are where client-facing incidents live. Agencies that map the surface before launch can scope retainers around the layers they actually cover, then charge for the ones they do not. A voice agent build illustrates the split: Cekura simulates thousands of personas and flags gibberish, interruption, and latency issues before go-live, while Hume AI layers emotion tagging and human rater feedback across 48+ emotions. Those are two different surface layers, two different line items. When a client asks why monitoring costs what it does, the answer is a coverage map, not a dashboard screenshot.
- Production Readiness GateConcept
The Production Readiness Gate treats evaluation as a contractual checkpoint rather than a post-launch cleanup task. Before any AI feature touches a client's live environment, it must clear a defined bar: traced agent behavior, scored response quality, and drift detection running on real traffic. Agencies that formalize this gate can price AI work as production-ready delivery instead of experimental builds, because the gate produces evidence the client can audit. The gate also caps downside: when an agent misbehaves, the trace log shows exactly which span failed and when, which shortens incident reviews from days to hours. A voice agent deployment illustrates the pattern well. Cekura simulates thousands of personas before go-live, then monitors live calls for gibberish, interruption, and latency signals, so the agency hands over a system with a documented pass record rather than a demo. Langfuse and Confident AI serve the same gate function for text and multi-model stacks.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- AI Evaluation Rule: Instrument Before You Scale Agent AutonomyEvaluation Rule
Wire tracing, scoring, and drift detection into the agent before you widen its autonomy or client exposure, not after the first failure.
- AI Evaluation Rule: Price the Eval Layer Into the Retainer Before the Second Agent ShipsEvaluation Rule
Bill evaluation and observability as a named retainer line from the first production agent onward, and treat any deployment without it as an unpriced liability rather than a completed deliverable.
- Evaluation Pipeline Before Launch vs Retrofit After Client EscalationDecision Framework
IF an agency is shipping LLM features or voice agents into a client retainer, THEN instrument tracing and scoring before the first production release, because failure modes surface as client-visible incidents rather than internal bugs. IF the agency has already launched and is fielding complaints, THEN treat the retrofit as a scoped remediation project with its own fee rather than absorbing it into existing delivery hours.
- Why AI Evaluation & Observability Stalls After the Pilot DemoFailure Pattern
- The Judge-Only Trap: Why AI Evaluation & Observability Collapses When Scoring Never Touches ProductionFailure Pattern
Delivery system
Blueprints and procedures for running it as a service.
- Production AI Readiness Audit and Eval Harness Build (10-15 days)Implementation Blueprint
A fixed-scope engagement that instruments a client's live LLM feature with tracing, scoring, and drift alerts, then hands over a scored baseline the client's team can defend in a board or procurement review. It converts an unmonitored AI deployment into a documented, retainer-ready production system.
- Pre-Launch Eval Gate (Onboarding)Operating Procedure
- Production Trace Triage (QA)Operating Procedure
- Client-Facing Eval Scorecard Handoff (Handoff)Operating Procedure
13 modules selected for Referee
Frequently Asked Questions
Answers about pricing, setup, implementation
Referee.chat runs multiple LLM models against a user-defined quality standard and judges the outputs based on evidence rather than model confidence. You set the bar (publication standard, formal proof, decision-grade, or custom), seat a panel of models, and Referee tells you which outputs meet the bar and why. For mathematical claims, it integrates with theorem.chat to formalize proofs in Lean against Mathlib.
The Free plan costs $1 one-time and includes every model and full transcripts, but matches are published publicly. The Credit plan costs $10 per month per seat (up to 6 seats), keeps matches private, and includes refundable unused credits and higher spending limits. You can also bring your own inference key and pay your LLM provider directly; Referee charges $0.05 per unit of infrastructure after the first $5.00 of usage, with no monthly seat fee.
Strategists and researchers compress quality-assurance cycles by running multi-model evaluations against publication or formal-proof standards. Project managers formalize AI-generated recommendations into decision-grade outputs with costed options and named risks. Content leads and operations teams use Referee as a quality gate to validate AI outputs before client delivery, replacing manual spot-checking with evidence-backed verdicts. Founders building internal AI workflows use Referee to lock in quality standards that persist across all future evaluations.
A strategist or researcher running 5+ AI evaluation cycles per week typically saves 3-5 hours per week by replacing manual multi-model comparison and quality review with a single Referee match. A project manager formalizing 2-3 AI-generated recommendations per week saves 2-3 hours by automating the evidence-gathering and risk-naming step. Payback depends on baseline QA time; teams spending less than 2 hours per week on AI review will see minimal ROI.
Referee integrates with 40+ model providers including Anthropic Claude, OpenAI GPT (including o1, o3, GPT-5 series), Google Gemini, Mistral, DeepSeek, Meta Llama, Cohere, AI21, Amazon Nova, and many others. You can also bring your own inference key and use Referee with your existing vendor relationships. The full model list is available on the Referee.chat platform.
Referee does not publish pricing or data-retention terms in the extracted content. Contact Referee.chat support for details on data export, deletion, and retention policies after cancellation.
Initial setup takes 30-60 minutes per team: define your quality standards (publication, formal proof, decision-grade, or custom), select your model panel, and run a test match. Rollout to the full team is low-friction; each team member can start running matches immediately after onboarding. No infrastructure changes or API integrations required.
Referee automates the evidence-gathering and verdict step of QA, but does not replace human judgment on whether the quality bar itself is appropriate for a given task. Use Referee to formalize and speed up QA workflows, not to eliminate human review entirely. Teams typically use Referee to flag outputs that fall short of the bar, then route those outputs back to the model panel or to human review.