Practice Lab
Real scenarios graded by AI — write prompts, build pipelines, debug agents. The AI grader scores every attempt.
live grading
Active scenarios
12
Avg AI grade
81%
Last 30 days
Completed today
3
Daily streak: 14 days
Grader model
sixfactors-grader-v3
live
Auto-rubric · multi-turn aware
Agents
In progress
Customer-support agent triage
- Resolution rate · empathy · escalation accuracy
- Turns: 6 / 12
- grader: You routed an angry billing case to tier-1 — escalation criteria not met.
Not yet graded
Resume
RAG
Graded
RAG retrieval bench — noisy corpus
- Recall@5 · MRR · latency budget · cost / 1k queries
- Turns: 8 / 8
- grader: Recall@5 = 0.81 (target 0.80). Latency p95 over budget — try smaller chunks.
- AI grade: 87
Resume
Safety
Needs rework
- Prompt injection defense
- Survives 12 known attacks · no system-prompt leak · no PII echo
- Turns: 5 / 10
- grader: 3 of 12 attacks succeeded. System prompt leaked via translation roleplay.
- AI grade: 42
Resume
Tool Use
Graded
Tool-using agent — flight search
- Correct tool selection · arg validation · graceful failure
- Turns: 9 / 9
- grader: Clean tool ladder. Add timeout handling on the booking step.
- AI grade: 94
Resume
Evals
In progress
Eval-set authoring drill
- Coverage · golden quality · regression sensitivity
- Turns: 3 / 6
- grader: Goldens skew positive — add 4 adversarial examples.
- Not yet graded
Resume
Voice
Available
Voice agent — appointment booking
- Turn-taking · interruption recovery · slot accuracy
- Turns: 0 / 14
- grader: Real-time grading — speak naturally, the AI scores latency and clarity.
- Not yet graded
Start
Invalid domain for site key.
ERROR for site owner:
Invalid domain for site key
reCAPTCHA