A game day is a planned failure exercise. The team picks a failure mode, induces it in a controlled way, and watches how the system and the responders react. Done well, it surfaces the gaps in monitoring, alerting, and runbooks before a real incident does.
This post is what a useful game day actually looks like, the patterns that make it productive instead of theatrical, and the common failures to avoid.
Why Game Days Exist
Most production incidents reveal that something the team thought was true was not. Alerts that "should fire" did not. Runbooks that "are documented" are out of date. Failovers that "definitely work" did not. The first time these gaps are discovered should not be at 3 AM during a real outage.
A game day is a low-stakes rehearsal. The failure is real (the system is actually broken in a controlled way), but the impact is bounded and the timing is chosen.
What a Game Day Is Not
- A drill — a drill walks through a scenario without inducing the failure. Useful, but a different exercise.
- Chaos engineering — automated, continuous, random failure injection. Game days are scheduled and scoped; chaos engineering is constant.
- A retrospective — a retro analyzes a past incident. A game day creates a controlled incident to learn from.
- A blameless postmortem — that comes after the game day, like any incident.
A game day is "let us deliberately break this thing during work hours and see what happens."
Scoping the Exercise
A useful game day is narrow. The whole exercise is over in a few hours, including setup and debrief.
Three good scoping patterns:
Single-component failure. "We will simulate a database failover at 2 PM Wednesday." One component, one well-understood failure, observable outcomes.
Single-runbook execution. "We will execute the runbook for 'primary cache unreachable' from start to finish, with the cache actually unreachable." Tests the runbook against reality.
Specific question. "Does our alerting actually catch a 50% latency degradation in the checkout flow?" Induce the condition, measure the response.
Bad scoping looks like "let us break everything and see." Too much surface, too many lessons, debrief becomes impossible.
The Three Roles
Successful game days have explicit roles:
Facilitator. Runs the exercise. Maintains the schedule, keeps participants on track, manages communications. Does not participate in the response.
Failure injector. Causes the failure at the agreed time. May be the facilitator, may be a separate role for complex failures.
Responders. The on-call team. Treat the game day as a real incident, with real urgency. The point is to expose how they actually respond.
Observers. Watch, do not participate, take notes for the debrief. Engineering managers, senior engineers, the team's tech lead.
Separating these roles prevents the failure injector from becoming the response coordinator, which is what happens when the same person runs everything.
The Schedule
A typical four-hour game day:
| Time | Activity |
|---|---|
| T-30 | Pre-brief: facilitator describes the exercise, sets ground rules |
| T-0 | Failure injection |
| T+0 to T+90 | Response (the main event) |
| T+90 | Restoration if needed |
| T+90 to T+120 | Debrief |
The pre-brief is essential. Responders should know what is happening at the meta level — they are responding to a game day failure, not a real one. Conflating real and rehearsal alerts is dangerous.
The debrief while memory is fresh is where the value comes from. Skip the debrief and you have just inflicted an outage on yourselves for no benefit.
Communication During the Exercise
Game day communications need a separate channel from production incidents.
- A
#gamedaySlack channel for the exercise - All game day messages prefixed
[GAMEDAY] - Status pages NOT updated
- Customer-facing communications NOT triggered
- Pages NOT escalated to people outside the exercise
Real production communication mechanisms should remain quiet. Confused customers and woken executives are not the lesson.
Failure Injection Patterns
What failures are worth injecting? Common, valuable ones:
- Single AZ/region loss for a multi-AZ system. The actual test of your HA design.
- Database failover from primary to replica. The runbook test.
- External API degradation. Slow responses, error spikes, full outage. Tests retry and fallback logic.
- Cache layer failure. What happens when Redis is down. Tests the load increase on the database.
- DNS resolution failure for an internal service. Tests fallback addressing.
- Disk full on a node. Tests monitoring and capacity alerts.
Choose failures the system should survive. A failure mode that takes down the whole system is not a game day; it is a chaos event with consequences.
Tools for Injection
For Kubernetes:
- kube-monkey, chaos-mesh, litmus — pod, network, and resource failures.
- Manual
kubectl delete podor scaling deployments to zero.
For AWS:
- AWS Fault Injection Simulator (FIS) — managed service for injecting failures.
- Manual stops of EC2 instances or RDS failovers.
For applications:
- Toxiproxy — controllable network proxy that can introduce latency, errors, drops.
- Application-level flags that simulate dependency failures.
Manual injection is fine for occasional game days. Specialized tools matter when you want to repeat the same scenario reliably.
The Debrief
The debrief is where the value emerges. It is not a postmortem — there is no incident to mortem. It is a guided conversation about what happened.
Useful questions:
- Did alerts fire? Did they fire in time?
- Did responders find the runbook? Did the runbook work?
- What did responders need that they did not have?
- Where did time get wasted?
- What was confusing?
- What surprised people?
The output is a list of follow-up work items: missing alerts, outdated runbooks, missing dashboards, unclear ownership.
The work items are the deliverable. A game day that produces no follow-up is either a perfect system (unlikely) or a poorly-debriefed exercise (more likely).
Cadence
Game days work best on a regular cadence:
- Monthly for a single team. One scenario per month, focused on what is most fragile or recently changed.
- Quarterly at the platform level. Larger scenarios involving multiple teams.
- Annually for major scenarios like region failover. The biggest exercises require the most preparation.
A team that has never run a game day should start with a small, low-stakes exercise. The first game day teaches the team how to run game days; subsequent ones produce more value.
Common Mistakes
- No pre-brief. Responders do not know it is a game day. They respond as if it is real, including customer comms. Damage.
- Real-channel communication. Game day alerts go to PagerDuty and wake real on-call. Confusion.
- No debrief. The exercise produces stress but no learning.
- Choosing failures that take the system down. A game day should test resilience, not destroy production.
- Same scenario every time. After two repetitions, the team is just executing a script. Vary the failure.
- No follow-up. The debrief produces a list of items that never get worked. The next game day re-discovers the same gaps.
When Game Days Are Not Right
- The team has no functional on-call yet. Build that first.
- The system has no monitoring or alerting. There is nothing to test.
- A real outage is in progress. Reschedule.
- The team is in crunch mode. Game days take time and energy.
A game day is an investment in resilience. It pays off when the team has the basics and wants to verify them, not when the basics are missing.
What Maturity Looks Like
A team that runs game days well:
- Runs them on a regular cadence without ceremony
- Has clear scenarios in a library, reused over time
- Generates and works through follow-up items
- Sees alert and runbook gaps shrink over time
- Treats real incidents as familiar, because they have rehearsed similar ones
The goal is not to never have outages. The goal is to be so well-prepared for the ones that happen that they are boring.
Building incident response practices for a team that has the basics but has not rehearsed them under pressure? We help teams design game days that produce real improvement, not theater. scopeforged.com