Stay Confident
Subscribe to our weekly newsletter to stay confident in the AI systems you build.

Introducing Jev-as-a-Judge: Faster, Cheaper, More Consistent Evals on Confident AI
Most teams sample 1% of production traffic for online evals and classification because every judge call costs money. Jev-as-a-Judge makes those decisions fast and cheap enough to run on 20%, 50%, or more of your traces, so every workflow on Confident AI has more data to work with.

Introducing Customer Monitoring: Product Analytics for B2B AI Agents
PMs should be able to understand how every account experiences their AI agent. Customer Monitoring groups users into the customers they belong to and shows each account's sentiment, issues, use cases, usage, and cost, all backed by traces and evals.

Introducing Category Correlation: Identify Which Production Signals Move Together
Filters only answer the combination you thought to ask. Category Correlation analyzes every pair of classifier labels across your incoming traces and threads, surfaces the relationships that matter, and links them back to the evidence.

Introducing AI Drift Detection: Find What Changed in Your Production Agents
Finding a production issue should not mean hours filtering traces. The Drift page automatically versions each agent configuration, surfaces anomalies and regressions, and shows exactly where each change is concentrated.

Introducing Voice AI Evals: Test the conversation, not just the transcript
Voice AI teams should not need one tool for audio testing and another for general AI evaluation and observability. Voice AI Evals bring both together on Confident AI, from persona-based testing to deterministic audio metrics.

Introducing confident-trace: OTel Native Production Monitoring for AI Agents
confident-trace is our new Python and TypeScript SDK for tracing production AI applications. Here's why we're separating it from DeepEval and what that means for your stack.
A Practical Guide to Building the Right Agent Evals
The right evals start with real production failures, not a list of generic metrics. Here is a practical workflow for discovering issues, validating them with humans, and turning them into durable evaluation coverage.

Introducing Report Templates: Build the report your team actually reads
Report Templates let you customize the reports Confident AI generates for your team. Build daily reports that dig into traces, identify where your AI agent is underperforming, summarize common usage patterns, and show the exact pages and sections you care about.

AI Agent Observability: Everything You Need to Know in 2026
Everything you need to know about AI agent observability in 2026 — traces, spans, and threads; online and offline evals; production monitoring; and closing the feedback loop so failures never repeat.

Introducing Synthetic Data Generation Pipelines: Customize how you generate data
Many teams already had great synthetic data generation pipelines running locally, but consolidating that work on one platform usually meant giving up flexibility. Synthetic Data Generation Pipelines bring that control into Confident AI: choose the sources to draw context from, wire them together, and tune each generation step.

