London

June 28–29, 2027

New York

September 15–16, 2026

Berlin

November 9–10, 2026

One shot at scale: Surviving the Super Bowl signup surge

This talk goes behind the scenes of how Fetch prepared for a massive Super Bowl traffic spike, scaling from around 1 signup per second to a target of 150,000. It explores the engineering decisions, architectural trade-offs, stress testing, and launch-day processes that helped the team manage risk when there was only one chance to get it right.

Speakers: Vadim Uchitel

Register or log in to access this video

Create an account to access our free engineering leadership content, free online events and to receive our weekly email newsletter. We will also keep you up to date with LeadDev events.

Register with google

We have linked your account and just need a few more details to complete your registration:

Terms and conditions

 

 

Enter your email address to reset your password.

 

A link has been emailed to you - check your inbox.



Don't have an account? Click here to register
September 29, 2026
NYC 27 Pre sale ticket block image

When Fetch aired its Super Bowl commercial, we had one chance to get it right. We ran a live $1.2M sweepstakes during the final two minutes of the game: anyone who opened the app entered automatically, and 120 winners were drawn and announced in-app in real time. Our normal signup load was ~1 RPS. Our target: 150,000. There was no rehearsal, and the sweepstakes carried legal obligations that meant certain operations had to be both fast and provably correct.

This talk focuses on the engineering leadership decisions that made this survivable. I’ll share how we decomposed the signup critical path into three buckets (must be synchronous, can be deferred, can be dropped) and how that framework shifted when a legally binding lottery meant some operations couldn’t be deferred or dropped. I’ll cover the architectural trade-offs: stripping the signup flow with feature flags, moving deferred work off the critical path with Kafka and SQS, and shifting verification to dedicated Redis clusters. Then I’ll explain how we decided which safety checks to temporarily relax, how we bounded that risk, and how we built confidence when our stress environment couldn’t replicate production.

Our stress testing caught some surprises early: Redis became CPU-bound at 150K RPS while DynamoDB held steady, the opposite of what we expected. But game day brought its own. Postgres connection limits quietly became a scaling ceiling, and an SMS rate limit we knew about still cost us valuable minutes of cross-team coordination during the live broadcast. Each surprise required real-time judgment calls, and our ability to respond came down to how we’d structured launch-day operations: defined roles, communication channels, and escalation paths designed with the same redundancy we’d built into the architecture.

If you’re preparing for a traffic spike with an immovable deadline, you’ll leave with a practical framework for decomposing a critical path, launch levers for controlling risk in real time, and hard-earned lessons about the gap between knowing a risk and being ready for it.

Key takeaways:

  • How to decompose a critical path into synchronous, deferred, and droppable work using feature flags and queues as scaling levers.
  • How stress testing at 150K RPS revealed counterintuitive bottlenecks (Redis, not DynamoDB, became the ceiling) and why preparing to be wrong mattered more than trying to be right.
  • How in-memory static assets and gradual rollout patterns can shield backend services when millions of users transition from a live event back into the regular app.
  • How to structure launch-day team operations (roles, comms, escalation) with the same redundancy you design into your architecture.