Week of August 15–21, 2026
Scheduled Collections 🔥
Set a Collection to run automatically (daily, weekly, or monthly) instead of remembering to trigger it yourself. Once it's scheduled, it just runs.
Collection pass-rate trends
Every Collection now tracks a pass-rate trend across its run history, both overall and broken out per suite. See whether quality is improving or slipping at a glance, instead of digging through individual runs to piece it together.
Run completion notifications
Set an email recipient on a Collection and get a summary the moment a run finishes: pass rate, per-suite breakdown, and a link to full results. No more checking back manually to see how it went.
Stop a run in progress
Suite and Collection runs now have a Stop button. Useful if you kicked one off by mistake, want to save on evaluation costs, or already know you don't need the rest. Anything already completed stays; only what hasn't run yet gets cancelled.
Faster runs, especially larger ones
Test suites and Collections now do independent work in parallel instead of one step at a time: multiple dataset rows, multiple suites in a Collection, and more all run concurrently wherever it's safe to do so. The bigger your suite or Collection, the more noticeable the speedup.
Clearer results when a check can't be evaluated
If a check genuinely can't complete (a timeout, for example), it's now distinguished from an actual failure instead of silently counting against your pass rate. You'll see it flagged separately, so a suite's score reflects what actually failed, not what couldn't be checked.
Reasoning alongside Intent verdicts
When a run is judged against your suite's Objective, the confidence score and the written reasoning now show together. Easier to see why a suite passed or failed, not just the number.
