Eval Studio
Write evals, golden sets, and regression tests for AI agents. Run them, see what breaks.
Summary
- Eval suites: 5 active
- Total goldens: 310
- Pass rate: 90.3% (+1.4)
- Avg runtime: 1m 41s
Running
/run
Regressions caught
- Today: 2
- Caught before merge
Eval Suites
Each suite is a versioned set of goldens. Pass requires all hard rubrics to clear and 0 regressions vs the last green run.
Suite Details
Agents
support-agent / golden / v3- Goldens: 119
- Pass: 4
- Fail: 1
- Last Run: 38s
RAG
rag-recall / chunk-strategy- Goldens: 65
- Pass: 12
- Fail: 3
- Last Run: 2m 12s
Safety
prompt-injection / red-team- Goldens: 26
- Pass: 5
- Fail: 1
- Last Run: 1m 04s
Tool Use
tool-call / arg-validation- Goldens: 56
- Pass: 0
- Fail: 0
- Last Run: 22s
Voice
voice-agent / latency budget- Goldens: 14
- Pass: 2
- Fail: 2
- Last Run: 4m 30s
Regressions caught today
Cases that were passing on the last green run and are now failing. The grader explains why.
rag-recall· long-table-mid-page regression- Chunk size raised from 800→1200 broke table boundaries
support-agent· angry-billing-tier-2 regression- Escalation threshold drift in v3 prompt
Error
Invalid domain for site key.