CATEGORY · 11 SERVICES

AI Evaluation & Observability Services

Compare 11 published AI Evaluation & Observability records on agency fit and pricing model. The category averages 5.0 of 10 and sits within Development, IT & Security.

11
Services
5.0
Avg score
11 of 11 match
11 services
AI serviceAI Evaluation Observability

Traccia

Traccia is an OpenTelemetry-native observability, evaluation, governance, and policy enforcement platform for AI agents.

Agency fit6.3 out of 10

Moderate fit

Best for AI development agencies

Review state
Published
Pricing
Tiered plans
Resale margin
49%
AI serviceAI Evaluation Observability

Langfuse

Langfuse is an open-source platform for tracing, evaluating, and monitoring LLM applications.

Agency fit5.8 out of 10

Narrow fit

Best for AI engineering teams

Review state
Published
Pricing
Tiered plans
Resale margin
57%
AI serviceAI Evaluation Observability

Agnost AI

Agnost AI continuously analyzes production AI agent conversations, finds where users get stuck, and turns high-impact patterns into reviewed fixes.

Agency fit5.7 out of 10

Narrow fit

Best for AI agent development teams

Review state
Published
Pricing
Free tier
Resale margin
57%
AI serviceAI Evaluation Observability

Cekura

Cekura provides automated QA, testing, and monitoring for voice and chat AI agents.

Agency fit5.7 out of 10

Narrow fit

Best for Conversational AI development teams

Review state
Published
Pricing
Plan + usage
Value score
3.3 out of 10
AI serviceAI Evaluation Observability

Confident AI

Standardize AI quality across teams with Confident AI.

Agency fit5.1 out of 10

Narrow fit

Best for AI development agencies

Review state
Published
Pricing
Free tier
Resale margin
49%
AI serviceAI Evaluation Observability

Referee

Referee.chat runs multiple AI models against your quality standards, with a referee that rules on evidence.

Agency fit4.9 out of 10

Narrow fit

Best for AI quality assurance agencies

Review state
Published
Pricing
Tiered plans
Value score
2.2 out of 10
AI serviceAI Evaluation Observability

Arize

Arize is the AI observability and evaluation platform for self-improving agents.

Agency fit4.8 out of 10

Narrow fit

Best for AI engineering teams

Review state
Published
Pricing
Tiered plans
Value score
3.1 out of 10
AI serviceAI Evaluation Observability

Hume AI

Hume AI provides data collection, evaluation, and human feedback infrastructure for voice and conversational AI systems.

Agency fit4.7 out of 10

Narrow fit

Best for Voice AI development teams

Review state
Published
Pricing
Free tier
Value score
3.1 out of 10
AI serviceAI Evaluation Observability

Braintrust

Braintrust is an AI observability and evaluation platform that helps teams monitor, test, and improve AI applications in production.

Agency fit4.6 out of 10

Narrow fit

Best for AI engineering teams

Review state
Published
Pricing
Plan + usage
Value score
3.1 out of 10
AI serviceAI Evaluation Observability

NEEDLE

NEEDLE is an open-source search benchmark that evaluates search APIs using agent-behavior queries across five verticals.

Agency fit3.9 out of 10

Weak fit

Best for AI agent development agencies

Review state
Published
Pricing
Pay per use
Value score
1.6 out of 10
AI serviceAI Evaluation Observability

Voker

Voker provides analytics and observability for AI agents.

Agency fit3.8 out of 10

Weak fit

Best for AI product teams

Review state
Published
Pricing
Free tier
Value score
3.1 out of 10

11 of 11 services shown

Showing 11 of 11 services

All 11 services shown