Launch Week 02 wrapped — explore all five launches

Product Changelogs

Released every Friday, once a week, every week — except for launch weeks.

August 7, 2026

Don't Cry Wolf

TGIF! Thank god it's features, here's what we shipped this week:

Not every ping deserves a page. Alerts now ship with priorities—critical, warning, error, and info—and integrations can filter by them, so Slack hears about the five-alarm fire while the polite FYI stays out of the channel. While we were sorting the inbox: MCP grew a fresh set of tools, report templates graduate from beta, red teaming picks up code vulnerability scanning, orgs can lock model providers to Bedrock-and-friends only, traces export as JSONL, and LLM spans finally log the endpoint they hit. Real wolves only. The boy who cried Slack is in timeout.

Changelog August 7, 2026

Added

  • Alert Priorities & Priority Filtering - Alerts now come in four flavors: critical, warning, error, and info. Set the severity when you create them, then filter integrations so Slack only hears about the five-alarm stuff while the rest of your stack still gets the full picture. Your pager stops treating every ping like the building is on fire. Cry wolf once, shame on you. Cry wolf every deploy—filter that.
  • More MCP Tools - Your coding agents can do more of the platform without leaving the chat. Spin up custom dashboards, annotate over MCP, kick off a risk assessment, or manage metric collections—plus governance assess and golden CRUD if you're feeling ambitious. Fewer tabs, more tools of the trade.
  • Code Vulnerability Scanning (Beta) - Red teaming can now scan for code vulnerabilities straight from the platform—invoke it, get findings, no DIY pipeline required. It's in beta and ready for you to poke holes in… well, your holes. Ship the attack, read the report. Code red, on demand.
  • Model Providers Policy - Organization admins can now disable model providers across every project—for example, Bedrock only, everywhere, no exceptions. One policy, zero "wait, who spun up that OpenAI key?" moments. Governance for the model menu. Provider locked.
  • LLM Span Endpoints - LLM spans can now log the endpoint they called, so traces show not just which model answered but which URL took the hit. Useful when you're routing through gateways, proxies, or three Bedrock regions and need to know which one actually ran. Follow the call to the end.
  • Trace Export as JSONL - Trace exports picked up JSONL as an additional format, so you can stream one JSON object per line into the tools that prefer that shape over a single blob. Same traces, line-by-line. JSON in a line.

Changed

  • Report Templates Out of Beta - Report templates graduated. What shipped as defaults last week is now a first-class, out-of-beta feature—build them, share them, and stop treating the template gallery like a science experiment. Officially not a drill. Report card: A+.

That's the drop for this week—see you next Friday.

July 31, 2026

Cost and Effect

TGIF! Thank god it's features, here's what we shipped this week:

The headliner: Project Cost Analysis, a Cost Insights tab in Data Usage that splits project spend across evals, signals, and platform features—then names the traces, spans, threads, and test runs quietly running up the bill. Everything else this week is about receipts too—audit logs now stream to Datadog and log admin actions, API keys got rotation and expiry APIs, GitLab joined the PR gate, and goldens went full CRUD. Cause, meet effect.

Changelog July 31, 2026

Added

  • Project Cost Analysis - Cost Insights came to Data Usage, and it brought receipts. Every project now splits its spend across online evals, offline evals, signals, and platform features, plots it over time by feature or by model, and shows what share of the org total you're personally responsible for. Then evaluation cost breaks down by where the evals actually ran—trace, span, thread, or test run—so the repeat offenders quietly re-incurring cost stop hiding in the aggregate. Turns out it's always the same three spans. Every token has a price on its head.
  • GitLab for PR Gate - The PR gate speaks GitLab now. Wire up merge requests the same way GitHub users have been gating theirs—quality checks run before the merge button does anything, and regressions get stopped at the door instead of in production. Same gate, new repo-tation.
  • Default Report Templates - Reports now come with templates out of the box, so you can generate something presentable without designing a document from a blank page first. Pick one, run it, send it. Off the shelf, on the money.
  • Golden CRUD Endpoints - Goldens went fully programmatic. New endpoints cover every CRUD operation, so you can create, read, update, and delete goldens straight from your pipelines instead of clicking through the dataset UI one row at a time. Automate the source of truth. Golden opportunity.
  • API Key Rotation, Grace Periods & Expiry - Key management grew up and got public APIs. Rotate keys programmatically, set a grace period so the old key keeps working while your services catch up, and give keys an expiry date so forgotten credentials stop living forever. Zero-downtime rotation, fully scripted. Everything in good key-ping.
  • Audit Logs to Datadog - Audit logs can now be exported straight to Datadog, so your compliance trail lands in the same place as the rest of your telemetry instead of sitting in a tab nobody opens. Ship the evidence where your eyes already are. Log and behold.
  • Admin Actions in Audit Logs - Admin actions are now part of the audit log too, which means the people with the most power to change things finally leave the clearest paper trail. Who did what, including the who that could do anything. Trail blazers.
  • Governance on Metric Data & Annotations - Governance runtime controls now understand the Metric Data and Annotations data models, extending your policies to two more places sensitive data actually lives. Fewer blind spots, same rulebook. Control freak, respectfully.

Changed

  • Classification Charts Move to Traces - The charts that used to hide in the classifications tab now render directly on the traces page, right next to the traces they describe. One page, one story, one less tab to remember. Chart where you are.
  • Fetch Trace Returns Everything - Fetching a trace now pulls the full trace, including data that had been offloaded to S3. No more partial payloads with the interesting parts missing because they were too big to stick around. The whole trace, every time. S3 what we did there.

That's the drop for this week—see you next Friday.

July 24, 2026

Every Version Everywhere All at Once

TGIF! Thank god it's features, here's what we shipped this week:

Your agent isn't one agent—it's dozens of variants living double lives in production. Automatic agent versioning now discovers every single one straight from your traces—no tags, no config, no archaeology. Add cost insights that name names, invites that finally exist, and filters you can actually share, and it's a big week in the agent-verse.

Changelog July 24, 2026

Added

  • Automatic Agent Versioning - We hash the LLMs and metadata in your traces to discover and version every deployment automatically, then let you compare dozens of variations on the fly. You ship, we catalog. Zero configuration, full receipts. A hash made in heaven.
  • Organization-Wide Cost Insights - One page, every project, all the spend. Finally find out which project has been eating the budget. (You know the one.) Follow the money.
  • Test Run Cost Insights - Test runs with traces now come with the bill attached. Quality on one axis, cost on the other, decisions suddenly defensible. Run the numbers.
  • Onboarding Invitations - Invite teammates during onboarding with token-based invites and an actual invitation accept page. Which, yes, we didn't have before. We're not proud either. Consider yourself invited.
  • Starter-to-Teams Upgrades - Upgrade from Starter to Teams on the fly, with your costs visible before you commit. No billing jumpscares. Dream big, Teams bigger.
  • Per-Metric Evaluation Models - Every metric can now pick its own judge. Match the model to the job, not the other way around. Here comes the judge.
  • Metric Collection Sample Rates - Evaluate a slice, not the firehose. Keep the signal, skip the bill. Please sample responsibly.
  • Shareable Filtered Views - Filters now live in the URL. Build a view, copy the link, and your teammate lands exactly where you wanted them. No more "okay now set these six filters and scroll down." URL in good hands.
  • Signals from Trace One - New projects turn on classifications and online evals automatically. Send traces, get signals. That's it. That's the setup. Love at first trace.

That's the drop for this week—see you next Friday.

July 17, 2026

Positive Developments

TGIF! Thank god it's features, here's what we shipped this week:

The headliner: classifier labels now have polarity, so when a signal trends up, you can finally tell whether to celebrate or panic. The outlook is decidedly positive—let's get into it.

Changelog July 17, 2026

Added

  • Classifier Label Polarity - Classifier labels can now be marked as positive, negative, or neutral, giving every trend the context it was missing. A spike in successful resolutions and a spike in safety violations may both point up, but your Signals can finally tell which direction is actually good. No more guessing whether up and to the right deserves applause or an alarm. Positive identification.
  • Async AI Connections - AI Connections no longer have to sit on the line while long-running agents finish the job. Asynchronous responses let agents keep working beyond a single request window, then return their results when they're ready—without timeouts cutting the conversation short. Good things come to agents that await.
  • Zero-Config Signals & Online Evals - New projects can go straight from sending traces to seeing Signals and online evals work out of the box. Sensible defaults are ready from the first trace, so there is nothing to configure before production quality starts showing up. Send first, set up never. Signal and deliver.
  • Cost Insights by Feature - Cost Insights can now break project spend down by feature, showing exactly which parts of your product are running up the tab. Follow the money from the project total to the feature responsible and optimize where it actually counts. Every token leaves a paper trace.
  • Annotation Date Filters - Annotations can now be filtered by date, making it easy to focus on a specific review window, compare periods, or find the label you know someone added last Tuesday. Your annotation history just became a lot less timeless. Date your data.
  • Metric Collection Sample Rates - Metric collections can now sample incoming traffic at the rate you choose, keeping online evaluation representative without evaluating every trace. Turn the volume down without losing the tune. Sample and hold.
  • Report Template Page Breaks - Report template sections can now start on a new page, giving every major section the clean entrance it deserves in exported reports. No more headings stranded at the bottom of the previous page. Time to turn over a new leaf.
  • Trace Detail Icons - Provider and integration icons now appear on trace details, so you can spot the services behind a span at a glance. Yes, icons made the changelog. They're small, they're obvious, and somehow we all survived without them until now. Tiny feature, big icon energy.

That's the drop for this week—see you next Friday.

July 10, 2026

Go With the Flow

TGIF! Thank god it's features, here's what we shipped this week:

The headliner: the new Flows page—in beta—a live map of how your agents call tools and models across every traced request, with failures, errors, and latency lighting up right where they cluster. Around it, onboarding can now scan your repo and open a tracing PR for you, custom skills teach your AI coding agents how your org actually works, MCP servers became first-class connections, online evals learned to sample traffic, and you can flag traces for review from the Observatory. Round it out with granular report emails, Hugging Face on evals, typed MCP vs. function tool calls, and a custom widget query endpoint. Let's get into it.

Changelog July 10, 2026

Added

  • Flows (Beta) - Meet the new Flows page: a live map of how your agents call tools and models across every traced request, with the trouble spots—failures, errors, latency—lighting up right where they cluster. Follow the paths your agents actually take and spot the mess before it becomes a mystery. It's in beta and ready for you to poke at. Flow state achieved.
  • Auto-Traced Onboarding - Setup just got a lot lazier, in the best way. Connect your GitHub repo during onboarding and Confident AI scans your codebase and opens a pull request that wires up tracing for you—no hunting through your app to hand-place spans. Review, merge, and you're already streaming traces. Instrument by pull request. Trace the easy way.
  • Custom Skills - Teach your AI coding agents how your org actually works. Author custom skills—plain-Markdown instructions with a description and body—at the org level or per governance policy, then install them into Cursor, Claude Code, or Codex so your agents onboard themselves to your conventions and compliance rules instead of guessing. Skills your agents actually follow. Skill issue, solved.
  • Flag Traces for Review - See a trace that needs a second pair of eyes? Flag it. Straight from the Observatory you can mark traces as requires review—one at a time or in bulk—then filter down to exactly the flagged pile when it's time to dig in, and unflag just as fast once it's handled. The messy ones stop hiding in the stream. Flag and drop.
  • Granular Emails & Report Emails - Email notifications grew a brain. Instead of blasting every project member with every ping, you now choose exactly who hears about what—test-run completions, alerts, and the brand-new report emails—recipient by recipient. Reports get their own email-only trigger that drops a fresh summary in inboxes the moment one is generated. Right message, right people, zero inbox riots. Mail it your way.
  • Online Eval Sampling - Online evals learned to pace themselves. Set a sample rate on a metric collection—or a project-wide trace and thread eval sample rate—so you score a representative slice of traffic instead of paying to grade every last request. High-volume projects keep the signal without the full bill. Sample smart, spend less. Rate yourself.
  • Hugging Face on Evals - Hugging Face joined the provider lineup. Point your evaluation model at a Hugging Face–hosted model and run LLM-as-a-judge on the open-weights option you actually want, instead of being fenced into the usual suspects—arena and model credentials speak Hugging Face too. Bring your own model to the party. Hug it out.
  • MCP Servers as Connections - MCP servers are now first-class citizens. Register one in project settings, authenticate it with OAuth client credentials or custom headers, pull in its available tools, and point evaluations and test runs straight at it. Your agents' tools finally have a home on the platform. Server's up.
  • MCP & Function Tool Calls - Tool calls now know what they are. Every tool call carries a type—Function or MCP—so traces, goldens, and test cases show at a glance whether your agent hit a local function or reached out over MCP, and tool-correctness evals can finally tell the two apart. Know your tools, judge them right. Call it like it is.
  • Custom Widget Endpoint - Dashboards went fully on-demand. The new widget query endpoint (POST /v1/widgets/query) lets you define a widget inline and pull its data back on the spot—quality, latency, cost, volume, any breakdown—without ever saving a dashboard, batching multiple lines in a single call at project or org scope. It now quietly powers every graph in the app, too. Query first, dashboard later. Widget your way.

That's the drop for this week—see you next Friday.

July 3, 2026

Significant Figures

TGIF! Thank god it's features, here's what we shipped this week:

The headliner: statistical significance for test runs—so when one run edges out another, you'll know whether it's a real improvement or just the sample size playing tricks on you. But the bigger story is how much of the platform went programmatic this week. Dashboards, Red Teaming, and Governance all grew full APIs, Jira joined the ticket-slinging integrations, and AI Connections learned to set themselves up. A lot that actually counts this week—let's get into the figures.

Changelog July 3, 2026

Added

  • Test Run Upgrades - Test runs got the full works this week. Statistical significance is now baked into every comparison, so when one run beats another you'll know whether it earned the win or just got lucky on a thin sample—no more reading tea leaves in a bar chart. Each test case can carry multiple generations instead of one lonely output, so you sample the model a few times before crowning a winner. And traces now attach to test runs straight from the API. Stop shipping on vibes, start shipping on proof. Significant upgrades.
  • AI Connections Upgrades - AI Connections picked up a stack of upgrades: query-parameter support, request logs so you can see exactly what went over the wire, and a new AI-powered flow that sets up your connection and debugs it for you when something's off. Wire it up, watch it work, let the AI fix the rest. Connect the dots.
  • Dashboards & Red Teaming APIs - Two of the platform's biggest surfaces went fully programmatic. The new Dashboards API covers full CRUD plus data access, so you can build, update, and pull dashboards straight from code, while the Red Teaming API lets you kick off assessment runs without ever opening the UI. Automate the reporting, automate the attacks. API-solutely everything.
  • Governance in the Admin SDK - Governance broke out of the console. confident-client now speaks fluent governance, so you can wire policies and controls straight into your own tooling and pipelines instead of clicking your way to compliance one checkbox at a time. Compliance-as-code, minus the clicking. Govern by code.
  • Jira Integration - Spot a bad trace, ship a Jira ticket. The new Jira integration joins GitHub and Linear on the "turn problems into tickets" bench, so the tool your team already plans in gets the memo automatically—no copy-paste pilgrimage required. Jira we go.
  • Thread Exports - Threads can now pack their bags and leave. Export full multi-turn conversations wherever the rest of your stack lives, so your production chats stop being trapped behind glass. Thread lightly, export freely.
  • Annotation Filters & Graphs - Annotations grew saved filters and graphs, so you can slice your labeled data down to exactly what matters and watch the trends surface instead of scrolling rows until your mouse wheel gives out. Label once, spot patterns forever. Note-worthy at a glance.
  • On-Prem Deployment Upgrades - Self-hosted deployments got a batch of enterprise-grade upgrades in one go: license-based feature gating so on-prem installs light up exactly the capabilities they're entitled to, plus internal root CA support for keeping connections locked down inside your own network. Everything the big deployments need, bundled up. License to ship.

Changed

  • Expanded Platform Model Support on Signals - Signals got more platform model support, now including a max tokens setting for tighter control over cost and output length when they fire. More knobs, same signal. Token of appreciation.
  • Better Dataset CSV Preview - The CSV preview got a serious upgrade—uploading a dataset now shows you a cleaner, clearer look at your data before you commit, so you catch the column that landed one cell to the left before it becomes next week's mystery bug. Same preview, much better look. Preview of coming attractions.
  • Real-Time Activity in Histograms - Histograms now update their activity in real time, so the bars move as the data lands instead of making you refresh to see what happened. Watch it fill in live. Histo-graphed as it happens.

That's a monster drop for one week—see you next Friday.

June 19, 2026

Govern Yourself Accordingly

TGIF! Thank god it's features, here's what we shipped this week:

The headliner: Governance grew into a full policy engine—spin up policies, stack them with controls, and roll them across every project in your org, with live compliance tracking baked in. Metrics now write their own criteria and rubrics. Classifiers picked up saved filters and workflow chaining that fires your whole pipeline off a single trigger. Round it out with the Gemini 3 series, charts that quit buffering, and paused-service heads-ups that actually tell you what to do about it. Let's get into it.

Changelog June 19, 2026

Added

  • Governance - Governance grew into a full policy engine. Spin up governance policies, stack them with pre-deployment and runtime controls, assign owners, and roll them out across every project in your org. A compliance matrix and daily-status views show who's passing and what needs action, the controls portfolio and project inventory lay out your whole estate at a glance, and audit logs keep the receipts. Compliance stopped living in a spreadsheet. Govern as you mean to go on.
  • AI Criteria & Rubric Generation - Metrics can now write their own criteria and rubrics with a little help from some very capable LLMs, so the blank-box-versus-the-word-"good" staring contest is officially called off. Dataset validation got sharper too, naming the exact required fields you're missing per metric instead of waving a vague yellow flag. Criteria met.
  • Classifier Filters & Workflow Chaining - Classifiers picked up saved filter configs and honest-to-goodness workflow chaining. Auto-classification eligibility now checks your default filters against trace snapshots, and the moment classification wraps, downstream queue and dataset ingestion kick off on their own. Pull one trigger, watch the whole pipeline fall in line. Chain reaction.
  • Gemini 3 Series Models - Gemini 3.1 Pro Preview, 3.5 Flash, and the 3-flash variants checked into the model catalog across evaluation and model selection. Fresh horsepower, no waiting list. Flash forward.

Changed

  • Per-Evaluation Model Overrides - Model separation landed: pick a model per evaluation for confidence-focused evals, backed by smarter platform- and feature-specific defaults across eval and generation flows. Point an unsupported provider at it and you get a clear error instead of a cryptic shrug. Model behavior.
  • Select All on Invitations - The org user multi-selector finally learned select-all and deselect-all, indeterminate header state and all. Inviting the whole team is one click now, not a finger workout. Select company.
  • Faster Metric Charts - Chart loading just got an upgrade. Metric charts that used to crawl on high-volume projects now snap into place, thanks to fresh caching under the hood. Chart-topping speed.

That's the drop for this week—see you next Friday.

June 12, 2026

Right on Time

TGIF! Thank god it's features, here's what we shipped this week:

This week, the platform learned to tell time. Tasks and alerts now respect onset, end-date, frequency, and run-once. Triage spread from traces all the way down to spans, threads, and test runs. Risk assessments picked up profile filters, official runs, and a BYOK option for Executive Insights.

Changelog June 12, 2026

Added

  • Flexible Scheduling for Tasks & Alerts - Dataset scheduling, exports, red teaming, and alerts all picked up the same scheduling kit—configurable onset, hard-stop on a specific run or specific date, flexible frequency, and a clean run-once switch. Your automations finally know when to clock in, when to clock out, when to repeat, and when to just punch in once and call it a day. Set your watch by them. Time well spent.
  • Triage on Spans, Threads & Test Runs - Last week, trace triage shipped to GitHub and Linear. This week, the same workflow rolls out to spans, threads, and test runs—any unit of investigation can now become a ticket in your team's issue tracker, not just full traces. Spot a problem, span a problem, ship a ticket. Span the gap between debugging and shipping.
  • Risk Profile Filters - Risk assessment views now filter by risk profile, so you can zero in on the threats that actually keep you up at night and scroll past the ones that don't. Focus the firepower, ignore the noise. Filter the threat, focus the fire.
  • Official Assessment Support for Risk Assessments - Risk assessments now carry an "official" badge, so your canonical red teaming runs stand apart from the scratch "let me just try one thing" sessions. Only the assessments that actually count make it into your risk history—everything else stays in the draft pile. Officially on the record.
  • BYOK for Executive Insights - Executive Insights reports can now be generated using your platform-preferred model. Bring your own key, bring your own model, and let your stakeholder reports come out in whichever flavor your stack already trusts. The c-suite finally gets the model you actually pay for. Key-note: yours.

Changed

  • AND/OR Toggle Across All Filters - Every filter on the platform now toggles between AND and OR logic. Stack conditions to narrow the haystack with surgical precision, or loosen the logic to widen the net—same filters, twice the modes, double the answers. And/or how you like it.
  • Model Credentials UI Polish - The model credentials flow got a polish pass—cleaner layout, fewer misclicks, less squinting at provider configs. Small surface, big quality-of-life upgrade. Credential check: passed.

Next week is Launch Week. Brace for launch.

June 5, 2026

On High Alert

TGIF! Thank god it's features, here's what we shipped this week:

The headliner is a brand-new Alerts page. We tore the old view down and rebuilt it so every alert keeps its full history, which means you find issues faster instead of squinting at Slack scrollback. The rest of the drop is no slouch either. Trace exports now let you pick a destination, official test runs keep dummy runs from crashing your eval history, and Salesforce and Snowflake join the knowledge base lineup for synthetic data. Risk assessment schedules graduated to the full attack engine, and a new org-level client in Python and TypeScript spins up projects on the fly. A lot to be alert about—scroll on.

Changelog June 5, 2026

Added

  • New Alerts Page - We tore down the old alerts view and rebuilt it from the ground up. Every alert now carries its full history, so you can see each one's complete track record—what fired, when, and how often—and drill from a noisy symptom to the actual root cause in a few clicks. Find issues faster, chase ghosts slower. Consider yourself alert-ed.
  • Trace Exports - Traces can now pack their bags and head wherever you need them. Export your traces and pick the destination, so your data lands exactly where the rest of your stack already lives—no more screenshot smuggling or copy-paste customs. Export control: yours.
  • Official Test Runs - Test runs can now be marked as official, which means your scratch runs, smoke tests, and "let me just try one thing" experiments stop crashing the party. Only the runs that count count, so your eval history finally tells the truth instead of a rumor. Make it official.
  • Salesforce & Snowflake as Knowledge Bases - Salesforce and Snowflake just joined the knowledge base lineup, so you can generate synthetic data straight from the systems where your real data already lives. Point, pull, and let the goldens write themselves—no exports, no glue code, no detours. Snow problem at all.
  • Attack Engine for Risk Assessment Schedules - Scheduled risk assessments now run the full attack engine instead of a watered-down sampler. Your recurring red teaming hits just as hard on autopilot as it does by hand, so threats get the full treatment whether you're watching or not. Set it, forget it, attack it.
  • Org-Level Client (Python & TypeScript) - A new client in both Python and TypeScript that speaks fluent org-level API key, built specifically for provisioning projects on the fly. Spin up new projects programmatically, at scale, with zero console clicking—just call it and watch a fresh project take the stage. The key to the kingdom.

May 29, 2026

We've Got an Issue

TGIF! Thank god it's features, here's what we shipped this week:

This week closes the loop from trace to ticket. GitHub and Linear integrations push problem traces straight into your issue tracker, the Integrations page got a card-based makeover with per-integration notification controls, and every alert that fires now leaves a full paper trail.

Changelog May 29, 2026

Added

  • GitHub & Linear Integrations - Spot a bad trace, ship a ticket. New GitHub and Linear integrations turn problem traces straight into issues in the tool your team already lives in—no copy-paste pilgrimage, no screenshot diplomacy, no "remind me which trace this was?" Trace, triage, ticket. Issue resolved.
  • Revamped Integrations Page - The Integrations page got a full card-based glow-up, with each integration getting its own card and its own fine-grained notification controls. Mute the chatty ones, crank up the critical ones, and stop drowning in pings that weren't yours to begin with. Card-carrying integrations, finally.
  • Alert History & Logs - Every alert that fires now leaves a paper trail. Browse a full history of triggered alerts—who got pinged, when, and why—so you can answer "wait, did that fire last Tuesday?" without spelunking through Slack scrollback. Alert and accounted for.
  • Traces on Red Teaming Test Cases - Red teaming test cases now come with full traces attached. When an attack lands a hit, you see exactly how the model got there—step by step, span by span—instead of squinting at the final output and reverse-engineering the path. Trace the threat, expose the route.
  • Dataset Version API Support - Dataset versioning is now fully scriptable via the API. Pin runs to specific dataset versions, automate version promotion, and keep your CI honest about which goldens it actually ran against. Version control: now actually under your control.

May 22, 2026

Queue Tip

TGIF! Thank god it's features, here's what we shipped this week:

Queues now know who to call, dashboards picked up every chart shape known to humankind, and traces went multimodal. Plus a stack of reliability fixes quietly landed underneath.

Changelog May 22, 2026

Added

  • Queue Assignment & Notifications - Annotation queues now route work to specific teammates and ping the assignee the second it lands. Take a number, get a name, get notified. Fewer "who's got this?" Slack threads, more "on it" replies. Queue the applause.
  • Provider & Integrations on Spans - Spans started naming names. Each one now tells you which provider and integration is actually doing the work, so you can stop pointing fingers and start pointing at the actual culprit—OpenAI, your vector DB, or that one piece of glue code you swore you'd refactor. Span-cific accountability, at last.
  • Multimodal Traces (PDFs & Images) - Traces are no longer text-only citizens. PDFs and images now ride shotgun through inputs and outputs alongside the words, so your multimodal model finally has a multimodal paper trail to match. Picture-perfect fidelity.
  • Risk Assessment Live Updates - Risk assessments stopped saving the drama for the season finale. Attack methods land and vulnerabilities surface live, so you can watch threats roll in as they happen instead of waiting for the credits. Live, laugh, threat-model.
  • Time-Series Tables & Every Graph Type - Dashboard widgets now sort themselves into time-series and categorical camps, joined by a brand-new time-series table and pretty much every chart shape known to humankind. If your data has a shape, we've already graphed it. Plot armor: equipped.
  • PDF & Image Export for Dashboards & Reports - Any dashboard widget or report now exports cleanly to PDF or image—deck-ready, doc-ready, leadership-ready. No screenshot diplomacy required. Export control: granted.
  • Thread Metadata Everywhere - Thread metadata is now stitched through the whole stack: ingestion picks it up, and the thread displayer, datatables, filters, and dashboards all read it back fluently. Tag once, slice forever, thread lightly.

Changed

  • Postgres Connection Pooling - We tracked down and squashed the connection pooling gremlin that occasionally turned Postgres into a waiting room of its own. The database is back to being a database, your requests are back to being responsive, and nobody has to ask "is it the DB?" first thing in the morning. Pooled resources, restored.
  • 2FA Is Back - Two-factor authentication returned from its brief stint in witness protection. Lock your accounts down properly again—with two factors, instead of two fingers crossed. Authenticate this.
  • The 95% Online Eval Error - The single error responsible for roughly 95% of online eval failures has been escorted off the premises, permanently. If your online evals were quietly losing runs to the void, the void is closed for business. Error-minated.
  • More Reliable Signals - Signals got a serious reliability pass under the hood. Fewer hiccups, less flakiness, exactly zero "is this thing on?" energy. Signal strength: restored, with bars to spare.

May 15, 2026

The Rules Have Changed

TGIF! Thank god it's features, here's what we shipped this week:

This week is about doing less. Online evals run themselves on rules you define in the UI, signals auto-classify into the issues actually showing up, and dataset reruns remember exactly how you set them up last time. Less wiring, more shipping.

Changelog May 15, 2026

Added

  • Evaluation Rules - Set up workflows to run online evals directly from the UI—no API call required. Pick your triggers, pick your metrics, pick your scope, and let the platform run the loop for you. Online evals used to be an API-only sport. Not anymore. Rule of thumb: less code, more coverage.
  • Prompt Editing in AI Connections - AI Connections now support prompt editing inside Arena and Experiments. Tweak prompts inline while you compare and iterate, without rebuilding the connection or leaving the page. Prompt and proper.
  • Evaluation Config History for Datasets - Every dataset run now saves its evaluation config to history. Rerun the same dataset later and bring back the exact same setup with one click. Reproducibility, but without the ritual. History doesn't have to repeat itself manually.
  • Auto-Classified Signals - Signals now auto-classify themselves into the issues actually surfacing across your traces. Find out what's wrong before you knew to look for it. Signal found, noise filtered.
  • Context & Retrieval Context for Multi-Turn Test Cases - Multi-turn test cases now support context and retrieval context fields. Test your RAG-powered conversations the same way you test single-turn outputs—same fields, more turns. Context collapse: averted.

May 8, 2026

Plot Twist

TGIF! Thank god it's features, here's what we shipped this week:

Welcome to Reliability Week. The plot has thickened—literally. Test Runs got a full analytics layer with heatmaps, bar graphs, and line-over-time charts that slice by any dimension you want (datasets, identifiers, hyperparams, models, prompts), so you can finally watch the trend instead of squinting at one run at a time. Offline Classification lets you classify traces and threads after the fact, and reclassify to backfill labels on data that came in before your rules existed. Auto-Surfaced Signals flips the question on its head—instead of you asking the data what's wrong, the platform tells you. Multi-Turn Evals leveled up across the board with variable interpolation, streaming prompts, and AI Connections support. And the views you actually live in—regression testing, thread displayer, test cases, Observatory tables—got a wave of polish.

Changelog May 8, 2026

Added

  • Advanced Test Run Analysis - The Test Runs page got an entire analytics layer. Aggregate every metric across every test run as a heatmap, bar graph, or line over time, and slice the view by any dimension that matters—datasets, identifiers, hyperparameters, models, prompts. Compare two slices side-by-side, toggle between Avg Score and Pass Rate, and watch the trend instead of squinting at a single run. Vibes are out, signal is in. Run the numbers.
  • Offline Classification - Classifiers now run offline. Classify traces and threads after the fact, and reclassify to backfill labels on data that came in before your rules existed (or got tagged wrong the first time around). Your old data finally caught up with your new rules. Classify later, sleep easier.
  • Auto-Surfaced Signals - Confident AI now auto-recommends signals on your traces, surfacing patterns, regressions, and weird-looking outliers without you needing to know what to look for. The dashboard tells you what's interesting, not the other way around. Signal acquired.
  • Multi-Turn Eval Upgrades - Multi-turn evals leveled up across the board: variable interpolation lets dynamic context, prior-turn references, and templated content play nicely across the whole conversation, and end-to-end support for streaming prompts and AI Connections means real-time conversations finally get real evals. No fake setup required. Turn up the volume.

Changed

  • Trace Comparison in Regression Testing - Regression testing now lets you diff traces, not just metric scores. When something regresses, see the actual trace-level difference instead of inferring it from a number that went down. Trace the regression.
  • Detail Displayer Upgrades - Both the Thread Displayer and Test Case Displayer got serious glow-ups this week. Component-level spans in threads are easier to scan and faster to navigate, and the Test Case Displayer got a polish pass that makes inspecting individual cases noticeably less squint-inducing. Cleaner hierarchy, faster context switching, fewer wrong clicks. Detail-oriented.
  • Revamped Test Cases Page - The Test Cases page picked up new tabs for end-to-end classification, component-level classification, and surfaced alignment insights. See exactly where each case lands across your eval pipeline at a glance, instead of clicking through three views to piece it together. Cases in point.
  • Sticky Column Headers in Observatory Tables - Column headers now stay pinned at the top of Observatory tables. Scroll to row 9,432 and still know which column is which. Stuck with you, in a good way.
  • Faster Test Case & Conversation Loading - Single-turn and multi-turn test cases now load dramatically faster, even on the gnarliest traces and longest conversations. Less waiting, more inspecting—on theme for Reliability Week. Load off your shoulders.

May 1, 2026

Health Check Yourself

TGIF! Thank god it's features, here's what we shipped this week:

This week is about knowing when things are healthy, knowing exactly how risky they are, and knowing your API keys cannot accidentally do too much damage. Health Dashboards give you a live pulse on evals, error rates, cost, and the signals that tell you whether your AI system is chilling or quietly catching fire. Comment Notifications keep the collaboration loop moving when someone tags you on the thing that needs attention. Customizable risk assessments, attack methods, and vulnerabilities let you shape red teaming around the threats your app actually cares about. And on the platform side, API keys and model credentials got a serious security glow-up: read-only keys, cleaner credential flows, org/project scoping, and suffixes that make keys easier to recognize before someone pastes the wrong secret into the wrong place. Prevention: still less annoying than incident response.

Changelog May 1, 2026

Added

  • Health Dashboards - Keep tabs on the health of your AI systems with dashboards for eval performance, error rates, cost, and the signals that tell you whether everything is fine or the model is doing interpretive dance in production. Less staring at charts hoping vibes improve, more knowing when to act. Health is wealth.
  • Comment Notifications - Comments now come with notifications, so tagged teammates actually see the thread, jump back into context, and help fix the thing instead of discovering it three standups later. Your comments have a pulse now. Notify and conquer.
  • Customizable Risk Assessments - Risk assessments are now fully customizable, including attack methods and vulnerabilities for custom evaluation steps. Test the risks that actually matter to your app instead of accepting a one-size-fits-all threat menu. Choose your own adventure, but make it adversarial.
  • Read-Only API Keys - Create API keys that can read but not write. Perfect for analytics, internal tooling, dashboards, and anything that should look around without touching the furniture. Least privilege just got easier to key into.
  • Model Credentials Flows - Model credential setup now has dedicated flows, making it easier to add, manage, and route provider credentials without turning setup into a scavenger hunt. Your models asked for better paperwork. We delivered. Credential where it's due.

Changed

  • Org- and Project-Scoped API Keys - API keys are now scoped to organizations or projects, with suffixes that make their scope easier to identify at a glance. Fewer mystery keys, fewer "wait, which environment is this?" moments, fewer self-inflicted footguns. Scope creep, but the good kind.
  • Auto-Formatted JSON in Dataset Goldens - JSON in dataset goldens now auto-formats on save. Your goldens stay readable, your diffs stay sane, and nobody has to pretend one-line JSON blobs build character. Format fortune favors the bold.

Next week is Reliability Week. Bring a helmet.

April 24, 2026

Better Work Is No Work

TGIF! Thank god it's features, here's what we shipped this week:

The best annotation work is the annotation work you never had to do. Auto-Annotate now takes the first pass across traces, spans, threads, and test cases, so your team can stop hand-labeling the obvious stuff and save human judgment for the weird, expensive, "why did the model say that?" moments. Multi-turn workflows got more automatic too: threads can become datasets with scenarios, ingestion tasks keep them fresh, and platform models can jump straight into simulations. Oh, and three beta stickers hit the floor this week: Code Execution, Queue Automations, and Dataset Workflows are officially stable. Less clicking. More knowing.

Changelog April 24, 2026

Added

  • Auto-Annotate Across Everything - Auto-Annotate now works on traces, spans, threads, and test cases. Let Confident AI take the first pass at labeling the chaos, then bring humans in where judgment actually matters. Less grunt work, more signal. Annotated for your convenience.
  • Thread Ingestion to Multi-Turn Datasets - Turn real user threads into multi-turn datasets, complete with scenarios. Your production conversations are no longer trapped in observability land—they can become eval fuel with a few clicks. From thread to test bed, no copy-paste pilgrimage required. Thread the needle.
  • Automated Thread Ingestion Tasks - Multi-turn datasets can now stay fresh automatically with thread ingestion tasks. Set the rules, let the pipeline run, and keep your evals fed with the kinds of conversations users are actually having. The dataset now has a metabolism. Ingest wisely.
  • Platform Models in Multi-Turn Simulations - Multi-turn simulations now support platform models. Bring the same model access you use across Confident AI into richer conversation testing, without detouring through yet another config maze. Simulations just got more well-modeled.
  • Prompts Tab in Multi-Turn Test Cases - Multi-turn test cases now have a dedicated Prompts tab, so you can inspect, edit, and understand the prompt behavior driving each conversation. Fewer mystery failures, fewer "where did that instruction come from?" moments. Promptly handled.

Changed

  • Arena Full-Screen Viewer - Arena now supports a full-screen viewer, because sometimes your model comparison deserves more than a cramped corner of the page. Go wide, judge harder. Arena seating upgraded.
  • Aggregated Turn Metadata in Arena - Arena now renders aggregated metadata for each turn, including tokens, latency, and cost. Compare outputs with the receipts attached, because vibes are useful but tokens still get billed. Meta made visible.
  • Confident Agent Pub-Sub Architecture - Confident Agent now supports a pub-sub architecture, making it more flexible for event-driven setups and distributed workflows. Your agent relay grew a nervous system. Published and subscribed.
  • Code Execution, Queue Automations & Dataset Workflows Are Stable - Code Execution, Queue Automations, and Dataset Workflows are out of beta and officially stable. The beta badges are gone, the features are staying, and your production workflows can stop side-eyeing the disclaimer. Stable geniuses.

April 17, 2026

@here Look At This Trace

TGIF! Thank god it's features, here's what we shipped this week:

Confident AI goes multi-player—and kills the context switch while it's at it. Comments are now live across traces, spans, threads, and test cases, and when someone @-mentions you, it lands in your Slack with a direct link back to the exact trace. No more "screenshot this span and DM it to me," no more five-tab scavenger hunts, no more "wait, which trace ID?" The conversation happens exactly where the data lives. That loop works because we also gave Slack & Discord a full glow-up this week—1-click setup, way more signals you can pipe through. And to the voice AI crowd: WebSocket response mode for AI Connections just shipped. We're coming for you. Custom Dashboards also picked up enough new widgets that the beta sticker is barely hanging on. Oh, and Claude Opus 4.7 is now available everywhere—Arena, Experiments, Evaluations, Platform. Plus Prompt Auto-Refinement on failing test cases, traces, and spans, and image support on annotations. Scroll down, there's a lot.

Changelog April 17, 2026

Added

  • Comments - Stop screenshotting spans into Slack DMs. Comments are now live on traces, spans, threads, and test cases—with full permissions and @-mentions that ping your teammate's Slack with a deep link straight back to the exact trace. No context switching, no "which trace again?", no losing the thread across three tabs. The conversation happens where the data lives. Oh, and you can mute or be muted. Finally, a proper comment section.
  • Revamped Slack & Discord Integrations - Our Slack and Discord integrations got a full rebuild: 1-click setup, way less config, and a lot more you can actually pipe through them—alerts, eval results, and @-mentions from comments, all landing in the channels your team already lives in. Channel your inner ops engineer.
  • WebSocket Response Mode for AI Connections - Voice AI, we're coming for you. AI Connections now speak WebSocket—true bidirectional, low-latency streaming for the stuff HTTP was never going to handle: voice agents, real-time assistants, long-running generations, anything where "wait for the full response" isn't an option. If you're building voice AI and you're not on Confident AI yet, this is your sign. Socket to 'em.
  • Metric FN/FP/TP/TN Over Time for Online Evals - Online Evals now plot false negatives, false positives, true positives, and true negatives over time. Catch metric drift before it catches you. Positively informative.
  • Native Annotation Test Cases - Annotations are now first-class test cases. Turn human feedback directly into evaluation data without any glue code or CSV gymnastics. Noted.
  • Tables & Big Number Widgets for Custom Dashboards - Two new widget types land in Custom Dashboards: Tables for row-by-row detail and Big Number for the one metric that matters most. Dashboards are inching closer to general availability—count on it.
  • Bar & Stacked Bar Graphs for Custom Dashboards - Bar and stacked bar charts join the Custom Dashboards widget lineup. Stack, compare, and break down your metrics any way you like. Raise the bar.
  • Prompt Auto-Refinement - Point at a failing test case (single-turn or multi-turn), trace, or span, and Confident AI will auto-refine the prompt for you—no more staring at a broken output and guessing which instruction to tweak. Your prompts, on autopilot. Refined to taste.
  • Image Support on Annotations - Annotations can now include images. Attach a screenshot of what went wrong, what it should've looked like, or the exact UI state that broke things. Human feedback with receipts. Picture perfect.
  • Claude Opus 4.7 Everywhere - Opus 4.7 is now available across Arena, Experiments, Evaluations, and the Platform. Pick your battles, pick your model. A true magnum opus.

Changed

  • Inline Table Editing - Editing values directly in tables got a serious polish pass—snappier, smarter, fewer misclicks, and a much better keyboard flow. The kind of upgrade you feel on every row.
  • PortKey Model Slug Fetching - Automatically fetch the model slugs available to your org's PortKey provider across Evaluation, Platform, and Arena. No more copy-pasting model names or guessing what's available. Slug it out no more.
  • Invitations for Organizations & Projects - Invitations now work at both the organization and project level. Bring people into the whole org or scope them to a single project—whichever fits the relationship. Invite-ing flexibility.

April 10, 2026

Back From Hiatus

TGIF! Thank god it's features, here's what we shipped this week:

Did you miss us? We missed you more, especially after last week's Launch Week! We're back with a loaded drop: Signals is in public beta—forget pre-defining metrics, Signals automatically surfaces issues, sentiment, and patterns across all incoming traces so you know what actually matters before you decide how to measure it. Confident Agent is live—a relay service that lets you expose internal endpoints to Confident AI without opening them to the public internet, so AI Connections just work with no security approvals or firewall hoops. Executive Reports enter public beta too: define your business KPIs and get daily generated reports against them. And for the org-level view: the Organization Governance Page lets you compare every project side by side on cost, metrics, annotations, and more.

Changelog April 10, 2026

Added

  • Signals (Public Beta) - Stop guessing which metrics to define upfront. Signals automatically detects issues, sentiment, and behavioral patterns across all incoming traces—so you discover what matters before you measure it. We're signaling a new era.
  • Confident Agent - A relay service that lets you expose internal endpoints to Confident AI via AI Connections—without opening them to the public internet. No more talking to security, no more firewall approval tickets. Just install the agent, point it at your endpoint, and Confident AI can reach it. Your security team can finally relax.
  • Executive Reports (Public Beta) - Define business-level KPIs and let Confident AI generate daily reports against them. Know exactly how your AI is performing in the language your stakeholders speak. Reporting for duty.
  • Organization Governance Page - See all your projects in one view and compare them head-to-head on cost, metrics, annotations, and more. Understand which projects are thriving and which need attention—across your entire org. Govern yourselves accordingly.

March 27, 2026

You Shall Not Merge!!!

TGIF! Thank god it's features, here's what we shipped this week:

The one you've been holding your breath for: Prompt Pull Requests & Approval Workflows are finally live—raise a PR on your prompt branch, let reviewers inspect diffs and eval results before signing off, and get a full audit trail of every change. AI Connections also got a major upgrade: a Postman-style layout, Auth0 and HMAC authorization, and direct trace linking to individual turns in multi-turn test runs. Plus: Thread Categorization with a configurable sample rate, and red teaming progress bars with more progress.

Changelog March 27, 2026

Added

  • Prompt Pull Requests & Approval Workflows - Raise a PR on any prompt branch. Reviewers see diffs and eval results side by side before approving, and every merge leaves a full audit trail of every change. Prompt engineering, meet version-control discipline. Approved.
  • AI Connection Authorization - AI Connections now support Auth0 SSO and HMAC signing. Secure your connections without the overhead. Consider it auth-orized.
  • Trace Linking to Turns in Multi-Turn Test Runs - AI Connections now link traces directly to individual turns within multi-turn test runs. Full visibility at every step of the conversation. The turn you've been waiting for.
  • Thread Categorization - Automatically categorize your threads to understand what your users are actually talking about. Set a sample rate to control how much traffic gets categorized. Categorically useful.

Changed

  • New AI Connection Layout - AI Connections get a Postman-inspired makeover: clean, familiar, and built for how you already think about API calls. Connect in style.
  • Improved Red Teaming Progress Bars - Progress bars for red teaming jobs got a polish pass—more granular, more informative, no more guessing how far along you are. Watch every step of your risk assessment unfold. Progress has definitely been made.

March 21, 2026

Branching Out

TGIF! Thank god it's features, here's what we shipped this week:

Buckle up—this is a big one. Prompt Branches bring proper version-control workflows to your prompts: branch, iterate, and merge without touching production. Custom Dashboards let you build your own Observatory views from scratch. Plus: OpenRouter and TrueFoundry are now available in Arena and Experiments, OpenInference tracing lands for Python and TypeScript, and enterprise auth gets a serious upgrade with HMAC & Auth0 support.

Changelog March 21, 2026

Added

  • Prompt Branches - Branch off your prompts, iterate safely, and merge back when you're ready. Your prompt engineering, with the same version-control discipline as your code. A real branch upgrade.
  • Custom Dashboards - Build your own Observatory dashboards from scratch. Pick your metrics, arrange your panels, tell your data's story. Your observatory, your _dash_board.
  • OpenRouter & TrueFoundry in Arena & Experiments - Two new model providers, one week. Access hundreds of models through OpenRouter or bring your fine-tuned TrueFoundry models—all available in Arena and Experiments. The route to more models just got shorter.
  • OpenInference Integration - Trace your LLM apps with OpenInference in both Python and TypeScript. Plug in, light up, see everything. Openly invited.
  • HMAC & Auth0 Support - Enterprise-grade authentication with HMAC signing and Auth0 SSO. Security that doesn't slow you down. Consider this auth-orized.
  • New Thread Displayer - Threads get a brand-new visual treatment—cleaner, faster, and easier to follow multi-turn conversations. Threads have never been so well-threaded.
  • AI Connections for Quick Runs & Experiments - Connect your AI provider directly for Quick Runs, and fine-tune temperature, top-p, and more right from the Arena and Experiments panel. No config files, no detours. Quick on the draw.
  • Error Bars in Observatory - Metrics now show confidence intervals so you know how much to trust the numbers. Finally, some margin for error.
  • Progress Bars for Risk Assessments - Red teaming jobs now show real-time progress instead of a spinner. Watch the risk assessment unfold. Progress has been made.

Changed

  • Transformers & Categories out of Beta - Battle-tested and production-ready. No more beta disclaimers—officially official.
  • User Analytics Upgrades - Total cost per user in the table, User ID filter on the Threads page, and click-through from Users to Traces. Your users, accounted for.
  • New Pagination & Arrow Navigation - Smoother pagination across the platform and arrow-key navigation for Spans and Threads. Keyboard warriors, we're turning the page for you.
  • Framework Deletion - You can now delete frameworks you no longer need. Sometimes you just need to let go.
  • General Stability & Performance Improvements - Bug fixes, reliability boosts, and the usual behind-the-scenes polish. The kind of changes you feel more than you see.

March 14, 2026

Version Control Freak

TGIF! Thank god it's features, here's what we shipped this week:

Datasets just got serious with Dataset Versioning—every change tracked, every version referenceable, no more "which dataset did we eval against?" Meanwhile, Replay Trace in Arena lets you re-run any production trace through Arena to compare models side-by-side on real traffic. And for the compliance-minded: Audit Logs are here.

Changelog March 14, 2026

Added

  • Dataset Versioning - Datasets now have full version history. Every edit, every addition tracked—so you always know exactly what you evaluated against. No more version of events that doesn't add up.
  • Replay Trace in Arena - Take any production trace and replay it in Arena. Compare how different models handle the same real-world input, side by side. It's the replay value you've been waiting for.
  • Audit Logs - Full visibility into who did what, and when. Every action logged, every change accounted for. Your compliance team just breathed a sigh of relief.

March 7, 2026

MC...What?!!

TGIF! Thank god it's features, here's what we shipped this week:

Headline first: Confident AI now has an MCP server (open-sourced on github)—plug your evals, datasets, and traces into any MCP-compatible client. Also shipping this week: automatic dataset curation from production traces, a wave of Observatory upgrades (custom column variable mapping, annotation tabs, category filters, metric columns), and PagerDuty for alerts.

Changelog March 7, 2026

Added

  • MCP Server - Plug Confident AI into any MCP-compatible client. Your evals, datasets, and traces—accessible from wherever you already work. The model context protocol is served.
  • Automatic Dataset Curation from Traces & Spans - The big one. Turn production traces and spans into curated datasets automatically. Your best (and worst) real-world examples, ready for eval—no manual curation required. Let your data curate itself.
  • Annotation Tabs in Observatory - Annotations now live in their own tabs, so you can flip between views without losing context. We're keeping tabs on your feedback.
  • PagerDuty Integration for Alerts - Route alerts straight to PagerDuty so the right people get paged at the right time. On-call never looked so connected.
  • Custom Column Variable Mapping - Map variables directly to custom columns in Observatory. Your data, your layout—no more squinting at mismatched fields. Finally, everything maps out.
  • Category Filters & Metric/Annotation Column Options - Filter by category and toggle metric or annotation columns on and off. Observatory now lets you see exactly what matters—no more, no less. Filter out the noise.

Changed

  • General Stability & Performance Improvements - Faster loads, fewer hiccups, smoother everything. The kind of changes you feel more than you see.

February 28, 2026

Prompt-ly Evaluated

TGIF! Thank god it's features, here's what we shipped this week:

Headline first: Prompt Evals are here. Think GitHub Actions, but for prompt commits and version releases—so every prompt change can trigger the checks that keep quality high and surprises low.

Changelog February 28, 2026

Added

  • Prompt Evals - The big one. Run evals on prompt commits and version releases automatically—CI for prompts, not vibes-based QA.
  • Support Ticket Submission Page - Need help? There's now a dedicated place to ask for it.
  • Trace Classification - Sort your traces into categories. Less chaos, more class.
  • Dataset Threads + Scenario Generation - Datasets now support threads, with scenario generation to spin up richer test cases.
  • Portkey Support in Arena and Experiments - Portkey now works in Arena and Experiments. The key to connected workflows.
  • SSE + HTTP Streaming for AI Connections - Stream responses over SSE or HTTP. Go with the flow.
  • Azure Key Vault Integration - Store secrets in Azure Key Vault. Your keys, under lock and cloud.
  • Org Settings Pages - New pages for roles, permissions, and API keys. Access control, finally under control.

Changed

  • Evaluate Buttons on Traces and Spans - Trigger evals directly from where issues appear, so troubleshooting is fewer clicks and more signal.

February 21, 2026

We Need to Talk. In Code.

TGIF! Thank god it's features, here's what we shipped this week:

Big week for the org-anized among us. Multi-turn evals go code-first, Vercel joins the family, and prompts finally get the observability they deserve.

Changelog February 21, 2026

Added

  • Code-Based Multi-Turn Evals - Introducing ConversationalTestCase for your codebase. All the power of multi-turn evaluation, now programmable. Time to have the talk with your chatbot—in code.
  • Vercel AI SDK Integration - Next.js devs, rejoice! Native integration with Vercel's AI SDK means you can trace and evaluate your ai package calls with zero friction. Ship fast, eval faster.
  • Transformers on Retrievers & Tools - Transformers aren't just for AI connection outputs anymore. Reshape retriever outputs and tool calls before evaluation. Your agentic RAG pipeline called—it wants its custom parsing back.
  • Organization-Wide Metrics - Define metrics at the org level and share them across all your teams. No more "wait, which faithfulness config are we using?" Standardize once, evaluate everywhere.

Changed

  • Prompt Observability - Track which prompts are running in production, when they were swapped, and how performance changed. Finally, prompt feedback on your prompts.

February 13, 2026

More Than Meets the AI

TGIF! Thank god it's features, here's what we shipped this week:

Transformers (Beta) are here and they're truly more than meets the AI. Reshape your traced data before evaluation—because not every trace deserves the full spotlight. Meanwhile, Prompt Studio just got a serious commit-ment upgrade with git-style versioning. Love is in the diff this Valentine's weekend.

Changelog February 13, 2026

Added

  • Transformers (Beta) - The biggest release this week, and it's more than meets the eye. Write custom code to transform your traced data—including individual spans—before evaluation. Don't want the whole trace? No problem. Cherry-pick exactly what matters.
  • Transformers on AI Connections - Got a JSON blob coming back from your model? Negative indexes on a list? Transformers let you parse and wrangle AI connection outputs however you need. Your data, your rules.
  • Prompt Commits - Every change to your prompt now creates a commit. Full history, no more guessing what changed or when. It's git log for your prompts, and it's beautiful.

Changed

  • Git-Based Prompt Studio - Prompt Studio is leaning hard into the git workflow. Commits, versions, diffs—everything you love about version control, now for your prompts. We're committing to this direction. (Pun intended.)

February 7, 2026

Let There Be Light (Mode)

TGIF! Thank god it's features, here's what we shipped this week:

Big week for visibility—both in your data and on your screen. We're launching 30+ additional Observatory graphs to surface insights, a Data Usage settings page for full transparency, and light mode is officially out of beta. Shine bright, friends.

Changelog February 7, 2026

Added

  • Data Usage Settings Page - Know thy data. A dedicated page to see exactly how your data is being used—because transparency isn't just a buzzword, it's a lifestyle.
  • Observatory Graphs - Finally, charts that slap. Visualize your observability data, spot trends before they spot you, and look like a genius in your next standup.
  • Code Evals (Beta) - G-Eval couldn't cut it? Write your own eval logic in code. We don't judge. Okay, technically we do—that's the whole point.
  • Multimodal Arena - Let your vision-language models duke it out. Two models enter, one model leaves with bragging rights.
  • AI Connection Upgrades - Tracing, list indexes key path, duplicate connections, max concurrency—the works. Your AI connections just got a glow-up.

Changed

  • Light Mode Out of Beta - Light mode is officially here to stay. Welcome to the bright side.
  • Faster Observatory Dashboards - We gave our dashboards a double espresso. Load times are now unreasonably fast.

January 30, 2026

Scaling New Heights

TGIF! Thank god it's features, here's what we shipped this week:

Welcome to our brand new changelog! We're kicking things off with better cost tracking, reliability improvements, and some serious scalability upgrades.

Changelog January 30, 2026

Added

  • Changelog - You're reading it! Subscribe to never miss a beat.
  • Custom Model Costs - Set custom cost-per-token for any model in your project settings. Finally, accurate cost tracking for fine-tuned and self-hosted models.
  • Request Timeout for AI Connections - Configure timeout limits for your LLM connections. No more hanging requests.
  • High-Volume Trace Ingestion - We've beefed up our trace handling with buffered ingestion. Traffic spikes? Bring 'em on.

Changed

  • Smoother Experiment Runs - Real-time evaluation progress is now more reliable with improved streaming.
  • Annotator Attribution - See who left that annotation. Credit where credit's due.
  • Faster Spans Loading - The spans tab now loads at lightning speed, even for trace-heavy projects.

January 23, 2026

Alert the Press, We're Going Multimodal

TGIF! Thank god it's features, here's what we shipped this week:

Big week! We're introducing alerts to keep you in the loop, shareable traces for collaboration, and multimodal support so your vision models don't feel left out.

Changelog January 23, 2026

Added

  • Public Trace Links - Share traces with anyone via a public link. Perfect for debugging with teammates or showing off to stakeholders.
  • Scheduled Alerts - Set thresholds, get notified. Never let a regression slip through unnoticed again.
  • Multimodal Evaluations - Images + text? We can evaluate that now. Test your vision-language models with confidence.
  • Evaluation Queue - Large eval jobs now queue up nicely instead of timing out. Go big or go home.

Changed

  • Snappier Dashboards - Graphs load faster. Like, noticeably faster. You're welcome.

January 16, 2026

On Cloud Nine

TGIF! Thank god it's features, here's what we shipped this week:

Azure fans, GCP enthusiasts—we see you. This week we're bringing the clouds to Confident AI so you can evaluate using your own infrastructure.

Changelog January 16, 2026

Added

  • Azure OpenAI Support - Connect your Azure deployment and run evals without leaving your cloud comfort zone.
  • GCP Vertex AI Integration - Drop in your service account key and you're off to the races with Google's models.
  • Top-K Filtering - Show me the top 10. Or bottom 5. Or whatever K your heart desires.

Changed

  • Faster Dashboards - We optimized the heck out of our aggregation layer. Graphs now load before you finish your sip of coffee.
  • Live Evaluation Progress - Watch your evals run in real-time with streaming progress updates. It's oddly satisfying.

January 9, 2026

Dashing Into the New Year

TGIF! Thank god it's features, here's what we shipped this week:

New year, new dashboards! We've redesigned how you visualize your LLM performance with customizable views and smarter breakdowns. And while you're at it, you can now take your security insights with you as PDF reports.

Changelog January 26, 2026

Added

  • Custom Dashboards - Build your own views. Save them. Make them yours. Finally, analytics that fit how you work.
  • Dimension Breakdowns - Slice and dice by model, environment, or any dimension. Compare apples to apples (or GPT-4 to Claude).
  • Risk Assessment Reports (PDF) - Generate custom risk assessment reports from your red teaming runs and download them as shareable PDFs. Perfect for reviews, audits, and internal security discussions. (Keep it confidential)

Changed

  • Fresh Dashboard Layout - Everything's been reorganized for better flow. Less clicking, more insights.
  • Readable Timestamps - Dates and times now look like actual dates and times. Revolutionary, we know.

January 2, 2026

I See What You Did There

TGIF! Thank god it's features, here's what we shipped this week:

Happy New Year! We're kicking off 2026 with a vision—literally. Multimodal evaluation is here, and your image-understanding models are about to get the testing they deserve.

Changelog January 2, 2026

Added

  • Multimodal Prompts - Drop images into your experiments. Test GPT-4V, Claude 3, Gemini, or whatever vision model you're building with.
  • Multimodal Test Cases - Build datasets with images + text. Because modern AI isn't just about words anymore.

December 26, 2025

Boxing Day Unboxing

TGIF! Thank god it's features, here's what we shipped this week:

Hope you had a great holiday! We kept it light this week, but still snuck in some dashboard goodies for you to unwrap.

Changelog December 26, 2025

Added

  • Duplicate Datasets - Clone any dataset with one click. Perfect for creating variations or backing up before big changes.
  • Better Invitation UX - Accepting team invitations is now smoother. New users get a clear onboarding flow instead of a confusing redirect.

Changed

  • Dashboard Reorganization - Dashboards now live under Home for easier navigation. One less click to your metrics.
  • Multiple Dashboards - Create different views for different needs. One for prod, one for staging, one for "what happened last night?"

December 19, 2025

Compare and Contrast

TGIF! Thank god it's features, here's what we shipped this week:

This week is all about perspective. New comparison features let you see how your models stack up—across time, segments, or whatever you want to measure.

Changelog December 19, 2025

Added

  • Custom AI Connection Payloads - Send custom parameters with your AI connections. Temperature, max tokens, stop sequences—whatever your model needs.
  • Comparison Mode - Put two time periods side-by-side. See exactly what changed and when. Debugging regressions just got easier.
  • Filter Presets - Save your favorite filter combos. One click to your most-used views.

December 12, 2025

A Metric Ton of Updates

TGIF! Thank god it's features, here's what we shipped this week:

We're laying the groundwork for multimodal evaluation and making threads easier to navigate. Plus, more control over your time ranges because "last 7 days" isn't always what you need.

Changelog December 12, 2025

Added

  • Multimodal Metrics - Purpose-built metrics for vision-language models. Evaluate what your eyes (well, your model's eyes) can see.
  • Enhanced Thread View - Follow conversations from start to finish. Every span, every trace, beautifully organized.

Changed

  • Flexible Time Ranges - Custom date windows are here. Pick any range. Go wild.

December 5, 2025

100 Ways to Evaluate

TGIF! Thank god it's features, here's what we shipped this week:

Why limit yourself to one LLM provider? With Portkey integration, you can now evaluate using 100+ providers. OpenAI, Anthropic, Cohere, local models—if Portkey supports it, so do we.

Changelog December 5, 2025

Added

  • Self-Served SSO - Set up SAML SSO for your organization without waiting on us. Enterprise security, self-service style.
  • Portkey Gateway - One integration, 100+ providers. Evaluate with whatever model you want, wherever it lives.
  • Custom Graphs - Build visualizations that actually match your KPIs. Your metrics, your way.

Changed

  • Optional Tags - Tags are now optional for evaluations. Less friction, faster setup. Just run the thing.

November 28, 2025

Permission Accomplished

TGIF! Thank god it's features, here's what we shipped this week:

This week we're locking things down—in a good way. Enhanced permissions give you granular control over who can do what, plus proxy support for enterprise networking needs.

Changelog November 28, 2025

Added

  • Multi-Factor Authentication - Add an extra layer of security to your account. Because passwords alone aren't enough anymore.
  • Role-Based Permissions - Fine-grained access controls are here. Admin, Editor, Viewer—you decide who gets the keys to what.
  • Proxy Agent Support - Running behind a corporate proxy? We got you. Configure your proxy settings and connect without hassle.
  • Observatory Filters - Filter your traces by any dimension. Find exactly what you're looking for, fast.

Changed

  • Smoother Authentication - Login flows are now more reliable, especially for enterprise SSO setups.

November 21, 2025

Lit-erally Amazing

TGIF! Thank god it's features, here's what we shipped this week:

LiteLLM joins the party! Evaluate with 100+ model providers through a single integration. Plus, better annotation tools to help you understand your data.

Added

  • LiteLLM Support - Connect your LiteLLM gateway and evaluate with OpenAI, Anthropic, Cohere, local models, and more. One config, endless possibilities.
  • Inherit Model Credentials - Projects can now inherit AI credentials from the organization level. Set it once, use it everywhere.
  • Annotation Upgrades - New annotation UI with better organization. Add notes, categorize results, and track quality over time.
  • Dataset Editor Improvements - Edit your golden datasets inline. No more export-edit-import dance.

Changed

  • Better Filter Persistence - Your filter settings now stick around between sessions. Less clicking, more analyzing.

November 14, 2025

Docker? I Hardly Know Her

TGIF! Thank god it's features, here's what we shipped this week:

Going on-prem? We've made self-hosting a breeze with proper Docker support. Plus, conversational evaluation just got smarter with multi-turn golden datasets.

Added

  • Docker Support - Full Dockerfile and docker-compose setup for on-premise deployments. Spin up Confident AI anywhere.
  • Conversational Goldens - Create multi-turn conversation datasets. Test your chatbots with realistic back-and-forth dialogues.
  • Auth On-Prem - Self-hosted deployments now support the same auth features as cloud. Enterprise-ready, wherever you run.

Changed

  • Simpler Environment Config - Streamlined environment variables for easier deployment. Less config, more evaluating.

November 7, 2025

Simulate to Dominate

TGIF! Thank god it's features, here's what we shipped this week:

Testing chatbots just got easier. Simulate entire conversations at scale and let your experiments run faster than ever.

Added

  • Conversation Simulations - Automatically simulate multi-turn conversations to test your agents at scale.

Changed

  • Faster Experiment Completion - Optimized the evaluation pipeline. Same results, less waiting.

October 31, 2025

Spooky Good Updates

TGIF! Thank god it's features, here's what we shipped this week:

It's Halloween and we've got treats, not tricks! Tool call streaming is here so you can watch your agents work in real-time, plus the prompt assistant just got a whole lot smarter.

Added

  • Tool Call Streaming - See your agent's tool calls as they happen. No more waiting for the final output to understand what went wrong.
  • Prompt Assistant Refactor - Our AI helper for prompt engineering got a major upgrade. Better suggestions, faster responses.
  • Max Concurrent Evaluations - Control how many evals run at once. Great for rate-limited APIs and not burning through your quota.

Changed

  • Better Error Messages - When things go wrong, you'll actually know why. Clearer errors across the platform.

October 24, 2025

Stream Team

TGIF! Thank god it's features, here's what we shipped this week:

Real-time just got more real. Streaming experiments mean you can watch evaluations progress live, and experiment edge cases are now handled gracefully.

Added

  • Streaming Experiments - Watch your evaluation progress in real-time. See results as they come in, not just when everything's done.
  • Multiple Winners Support - Experiments can now have multiple winners (or no winner). Because sometimes it's a tie.

Changed

  • Experiment Results Handling - Better handling of edge cases when experiments have unusual outcomes.
  • Runtime Environment Config - Cleaner configuration for different deployment environments.

October 17, 2025

Dataset Your Sights Higher

TGIF! Thank god it's features, here's what we shipped this week:

Datasets just got an upgrade. Evaluate prompts directly from datasets and manage your LLM connections with a dedicated page.

Added

  • Dataset Evaluate Prompts - Run evaluations directly from your datasets. Select a prompt, pick your metrics, and go.
  • AI Connection Page - A dedicated place to manage all your AI connections. See what's configured, test connections, and troubleshoot.
  • Prompt Version Labels - Tag your prompt versions with custom labels. Production, staging, experimental—organize however you want.

Changed

  • Smarter Project Fetcher - Projects load faster and more reliably, especially for organizations with many projects.

October 10, 2025

Gemini Rising

TGIF! Thank god it's features, here's what we shipped this week:

New model support and Arena improvements! Gemini joins the party, and the Arena just got more flexible with variable mapping.

Added

  • Gemini Support - Google's Gemini models are now available for evaluations. Test your prompts against the latest from Google.
  • Arena Variable Mapping - Map your test case variables to prompt placeholders. Makes testing prompt variations a breeze.
  • Arena Endpoints - New API endpoints for programmatic Arena access. Automate your prompt testing workflows.

Changed

  • Improved Model Selection - Cleaner UI for picking which model to use. All your options, clearly organized.
  • Better Trace Display - Trace details now render more cleanly, especially for long outputs.
Built byConfident AI