Skip to content

Proving it works

Grade a change before a caller ever hears it.

Build a dataset of real scenarios and the criteria a good answer should meet, then run a candidate prompt against it and watch a live call while it happens.

How it works

What the capability actually does.

Test cases from real questions

A dataset holds the scenarios worth checking and the criteria a good response should meet, built from questions people actually ask rather than ones nobody asks.

Run a candidate prompt before it ships

Point a proposed system prompt at the dataset and run it. The result is graded against your own criteria, not a subjective read of one transcript, and the candidate runs on the same model class as a real task rather than a cheaper stand-in.

Orchestrators get the same treatment

The same dataset, run, and comparison flow applies to an Orchestrator, where the candidate is executed as a real agent tree rather than a single generation.

Watch a call while it happens

Live call monitoring streams a conversation in progress rather than only its transcript afterward, for the period a new agent is still earning trust.

In the product

The screens this happens on.

An eval dataset with a test case, a candidate system prompt, and a completed run graded against it

A prompt change is a hypothesis

A dataset holds the scenarios worth checking and the criteria a good answer should meet. A candidate prompt runs against all of them and is graded against those criteria, on the same model class a real task would use, so the result reflects what would actually happen rather than a cheaper approximation of it.

  • The same flow covers Orchestrators, where the candidate runs as a real agent tree
Boundaries

What it deliberately does not do.

Stated here rather than discovered in production. Every one of these is a real constraint, and knowing them before you buy is worth more than a longer feature list.
  • Evals and live call monitoring are off on Demo and Basic. The 14 day trial unlocks both temporarily; a permanent licence needs Pro or Enterprise.
  • An eval grades what a candidate prompt actually produced against your dataset. It does not predict how a change will perform on a case nobody wrote.
  • Grading is a model judging a model. It is a great deal better than reading one transcript and guessing, and it is not a proof.