Free course

Agent Evals & Safety: Trust but Verify

Evaluate and harden AI agents: read failure traces, write deterministic checks over trace JSON, and defend against prompt injection — graded on the fix.

AdvancedAgentic AICertificate
6 modules 32 lessons 11 enrolled ~10h of material
Agent Evals & Safety: Trust but Verify

By the end

What you'll build

  • Classify agent failures with a structured taxonomy, including the silent 'corrupt success' mode.
  • Read and debug nested agent traces — thoughts, tool calls, observations, spans — to locate the exact point a run went wrong.
  • Bisect a run to separate root cause from downstream symptom, and recognize healthy recovery that only looks like failure.
  • Write deterministic assertion checks over trace JSON that grade final state, not prose.
  • Choose between deterministic checks and LLM-as-judge, accounting for judge biases like position and self-enhancement.
  • Compute and interpret agent eval metrics: pass@1, pass@k, pass^k reliability, and flakiness.
  • Model the prompt-injection threat (direct, indirect, tool-output) per OWASP LLM01:2025 and design layered defenses.
  • Apply defensive patterns: data/instruction separation, least-privilege tools, allow-lists, human-in-the-loop, and read-back verification.
  • Practice defensive red-team thinking against sandboxed agents — graded on the fix, never the attack.
  • Build a reusable eval suite and run it in a real open-source harness on your own machine.

The shape of it

How this course works

Short lessons

32 lessons across 6 modules, each small enough to finish in one sitting.

Practice as you go

Every lesson ends with a small space for what you noticed — the doing is the learning.

A certificate at the end

Finish the course and earn a certificate anyone can verify with a link.

Ready when you are.

Make an account and this course opens up — your progress is saved from the very first lesson.