London

June 28–29, 2027

New York

September 15–16, 2026

Berlin

November 9–10, 2026

Ms. Frizzle would be a great SRE

Five frameworks from a century of education research that map directly to eval design for AI agents, and why building the eval system before the product changes everything.

Speakers: Andrew Zigler

Register or log in to access this video

Create an account to access our free engineering leadership content, free online events and to receive our weekly email newsletter. We will also keep you up to date with LeadDev events.

Register with google

We have linked your account and just need a few more details to complete your registration:

Terms and conditions

 

 

Enter your email address to reset your password.

 

A link has been emailed to you - check your inbox.



Don't have an account? Click here to register
September 29, 2026
NYC 27 Pre sale ticket block image

Education researchers have spent a century figuring out how to build environments where complex, unpredictable systems self-correct through structured feedback. Software engineers building AI agents are solving the exact same problem from scratch and ignoring all of it.

This talk bridges that gap. I’m a former classroom teacher turned AI engineer, and I’ll walk you through five pedagogical frameworks that map directly to eval design patterns for AI agents: backward design (define success criteria before you build), formative assessment (eval continuously, not just at the end), rubric design (multi-dimensional scoring instead of pass/fail), error analysis (categorize failure modes because same symptom doesn’t mean same cause), and differentiated feedback (the agent, the user, and the knowledge base each need their own signal channel).

What ties them together is back pressure: each framework is a way to capture signal from problems and route it to where it drives change. That’s what makes a system self-correcting instead of just self-reporting.

Most engineering teams building AI agents design the system first and bolt on evaluation after. I think that’s backwards. The eval system should come first because it defines what “correct” looks like, and once you have that, you’ve given the system the ability to feel its own failures and iterate on them.

I’ve applied this methodology to production RAG agents and AI hackathons alike. The core insight is simple: the best eval systems aren’t tests, they’re environments. And nobody knows more about designing those environments than teachers.

This talk gives engineering leaders a concrete framework for building AI systems that steer themselves through structured feedback, drawn from a discipline most of us left behind after graduation.

Key takeaways

  • Backward design from education theory gives you a concrete methodology for eval-first AI architecture: define success criteria before building the system.
  • Pass/fail evaluation destroys information; multi-dimensional rubric design from pedagogy produces the granular signal agents need to self-correct.
  • Back pressure is the connecting principle: every pedagogical framework in this talk is a mechanism for capturing signal from problems and routing it upstream before those problems harden into production incidents.
  • AI agent interactions have three participants (agent, user, knowledge base) and each needs its own feedback channel; most teams only instrument the agent.
  • The eval system isn’t a test suite you bolt on at the end; it’s the steering wheel, and it should be the first thing you build.