Evaluating Agents
Test non-deterministic, multi-step agents the way you'd test any critical system: define what success means, build a suite of representative task-based evals with clear pass criteria, use judges (including LLM-as-judge, carefully) to score open-ended outputs, and run the suite on every change so you catch regressions before your users do — turning 'it seemed to work' into evidence.
Sign in to start your 85-day journey
This public overview is open to everyone. Members get the complete lesson, hands-on tasks, media, questions, feedback, and saved progress.