## Practice Lab

Real scenarios graded by AI — write prompts, build pipelines, debug agents. The AI grader scores every attempt.

### live grading

**Active scenarios**

12

**Avg AI grade**

81%

**Last 30 days**

Completed today

3

**Daily streak:** 14 days

**Grader model**

sixfactors-grader-v3

**live**

Auto-rubric · multi-turn aware

### Agents

**In progress**

Customer-support agent triage

- Resolution rate · empathy · escalation accuracy
- Turns: 6 / 12
- grader: You routed an angry billing case to tier-1 — escalation criteria not met.

**Not yet graded**  
Resume

**RAG**

**Graded**

RAG retrieval bench — noisy corpus

- Recall@5 · MRR · latency budget · cost / 1k queries
- Turns: 8 / 8
- grader: Recall@5 = 0.81 (target 0.80). Latency p95 over budget — try smaller chunks.
- AI grade: 87

**Resume**

**Safety**

Needs rework

- Prompt injection defense
- Survives 12 known attacks · no system-prompt leak · no PII echo
- Turns: 5 / 10
- grader: 3 of 12 attacks succeeded. System prompt leaked via translation roleplay.
- AI grade: 42

**Resume**

**Tool Use**

**Graded**

Tool-using agent — flight search

- Correct tool selection · arg validation · graceful failure
- Turns: 9 / 9
- grader: Clean tool ladder. Add timeout handling on the booking step.
- AI grade: 94

**Resume**

**Evals**

**In progress**

Eval-set authoring drill

- Coverage · golden quality · regression sensitivity
- Turns: 3 / 6
- grader: Goldens skew positive — add 4 adversarial examples.
- **Not yet graded**  
Resume

**Voice**

**Available**

Voice agent — appointment booking

- Turn-taking · interruption recovery · slot accuracy
- Turns: 0 / 14
- grader: Real-time grading — speak naturally, the AI scores latency and clarity.
- **Not yet graded**  
Start

**Invalid domain for site key.**

ERROR for site owner:

Invalid domain for site key

reCAPTCHA
