Register or log in to access this video

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
As large language models move from research to production, engineering leaders are facing a new kind of challenge: how do you test, deploy, and govern systems that can behave unpredictably? At GitHub, we found that reliability wasn’t enough — trust became the core metric. Building our internal coding agents required rethinking how we validated AI output, handled uncertainty, and collaborated across disciplines.
This talk explores that journey: how we built testing pipelines for non-deterministic systems, defined “ethical success criteria,” and aligned engineers, product managers, and legal teams around shared principles for responsible AI. I’ll share the technical patterns that worked — like controlled prompt experiments, sandboxed agent evaluation, and feedback loops — as well as the cultural lessons from introducing ethical review processes into fast-moving engineering work. Attendees will leave with concrete tools for building trustworthy AI systems, a playbook for leading cross-functional conversations about safety and ethics, and practical insights for balancing innovation with accountability.
In a world where AI capabilities evolve faster than our processes, this story is a reminder that trust is built, not assumed — and that engineering leaders have a critical role in making it real.
Key takeaways:
- Learn how to test and evaluate non-deterministic LLM behavior in production systems.
- See frameworks for defining and measuring “ethical success” beyond accuracy.
- Understand how to align engineers, PMs, and legal on responsible AI principles.
- Apply lessons for balancing speed, experimentation, and accountability in AI projects.