Practice Lab

Real scenarios graded by AI — write prompts, build pipelines, debug agents. The AI grader scores every attempt.

live grading

Active scenarios

12

Avg AI grade

81%

Last 30 days

Completed today

3

Daily streak: 14 days

Grader model

sixfactors-grader-v3

live

Auto-rubric · multi-turn aware

Agents

In progress

Customer-support agent triage

  • Resolution rate · empathy · escalation accuracy
  • Turns: 6 / 12
  • grader: You routed an angry billing case to tier-1 — escalation criteria not met.

Not yet graded
Resume

RAG

Graded

RAG retrieval bench — noisy corpus

  • Recall@5 · MRR · latency budget · cost / 1k queries
  • Turns: 8 / 8
  • grader: Recall@5 = 0.81 (target 0.80). Latency p95 over budget — try smaller chunks.
  • AI grade: 87

Resume

Safety

Needs rework

  • Prompt injection defense
  • Survives 12 known attacks · no system-prompt leak · no PII echo
  • Turns: 5 / 10
  • grader: 3 of 12 attacks succeeded. System prompt leaked via translation roleplay.
  • AI grade: 42

Resume

Tool Use

Graded

Tool-using agent — flight search

  • Correct tool selection · arg validation · graceful failure
  • Turns: 9 / 9
  • grader: Clean tool ladder. Add timeout handling on the booking step.
  • AI grade: 94

Resume

Evals

In progress

Eval-set authoring drill

  • Coverage · golden quality · regression sensitivity
  • Turns: 3 / 6
  • grader: Goldens skew positive — add 4 adversarial examples.
  • Not yet graded
    Resume

Voice

Available

Voice agent — appointment booking

  • Turn-taking · interruption recovery · slot accuracy
  • Turns: 0 / 14
  • grader: Real-time grading — speak naturally, the AI scores latency and clarity.
  • Not yet graded
    Start

Invalid domain for site key.

ERROR for site owner:

Invalid domain for site key

reCAPTCHA