Runbook Automation: Turning Wiki Pages into Executable Code

Philip Rehberger Oct 3, 2026 7 min read

Move from "screenshot the steps" to runbooks the on-call can run with one command.

A runbook is documentation an engineer reads at 2 AM to fix a problem they did not cause. The traditional runbook is a wiki page with a list of steps. The automated runbook is a script the engineer runs with one command. The difference matters at 2 AM.

This post is what runbook automation actually looks like, the patterns that hold up, and the cases where automation is the wrong answer.

The Problem With Wiki Runbooks

The traditional runbook lives in Confluence, Notion, or a markdown file. Steps look like:

  1. SSH to the primary database
  2. Run psql -U postgres -c "SELECT pg_promote_replica('replica-1')"
  3. Update the application config to point at replica-1
  4. Restart the application servers
  5. Verify the new primary by querying ...

This works, sort of. The problems show up under pressure:

  • Outdated steps. Half the time the steps are wrong because the system changed and nobody updated the wiki.
  • Copy-paste errors. Step 2's SQL has a typo that nobody caught.
  • Missing context. "Restart the application servers" is one line; in practice, this is twelve commands across three hosts.
  • Manual repetition. Each runbook execution is a fresh chance to make a mistake.

When you have a runbook executed dozens of times per year, the cost of these failures adds up.

What Automation Looks Like

The automated version expresses the steps as code:

#!/usr/bin/env bash
# Runbook: promote_replica_to_primary
# Usage: ./runbooks/promote_replica_to_primary.sh <replica-id>

set -euo pipefail

REPLICA_ID="${1:?usage: $0 <replica-id>}"

echo "Confirming replica is healthy..."
./scripts/check_replica_health.sh "$REPLICA_ID" || exit 1

echo "Promoting replica $REPLICA_ID to primary..."
aws rds promote-read-replica \
  --db-instance-identifier "$REPLICA_ID"

echo "Waiting for promotion to complete..."
aws rds wait db-instance-available \
  --db-instance-identifier "$REPLICA_ID"

echo "Updating application config..."
./scripts/update_db_endpoint.sh "$REPLICA_ID"

echo "Restarting application servers..."
./scripts/rolling_restart.sh app-servers

echo "Verifying new primary..."
./scripts/verify_primary.sh "$REPLICA_ID"

echo "Done. New primary: $REPLICA_ID"

One command. Same result every time. The runbook is in git; it gets reviewed, tested, and updated like any other code.

The wiki page now describes when to run the script, not how to do the steps.

The Levels of Automation

Runbook automation is a spectrum, not a binary.

Level 0: Wiki only. Documented steps, manually executed. Where most teams start.

Level 1: Scripted. Individual steps are scripts. Runbook is "run these scripts in this order."

Level 2: Composed. A single script runs the whole runbook. One command, many steps.

Level 3: Self-service. The script is exposed via a UI, ChatOps, or button. Non-engineers can run it.

Level 4: Auto-remediation. The system detects the condition and runs the script automatically. Human is notified but does not have to act.

Most teams should target Level 2 for common runbooks. Levels 3 and 4 are for specific scenarios where the cost of human latency is high.

When Auto-Remediation Is Right

Auto-remediation — Level 4 — is tempting but dangerous. It only works when:

  • The detection is reliable (very low false positives)
  • The action is safe (cannot make things worse)
  • The action is reversible (can undo if wrong)
  • The cause is well-understood (you trust the system to handle it)

Examples that fit:

  • Auto-scaling adding capacity when CPU saturates
  • Restarting a process that has hit a known memory leak threshold
  • Failing over from an unresponsive cache node

Examples that do not:

  • Database failover (you want a human to confirm)
  • Data correction (mistakes are expensive)
  • Security responses (false positives lock out users)

A good test: would you let this run at 3 AM with no human review? If not, do not auto-remediate.

The ChatOps Pattern

A common Level 3 implementation: runbooks exposed as Slack commands.

/runbook promote-replica replica-1

The Slack app validates inputs, checks permissions, executes the script, and reports results in the channel. Visibility, audit trail, and self-service in one mechanism.

The key design choices:

  • Permissions enforced before the script runs (not everyone can promote a database)
  • The action is logged with who ran it and when
  • Sensitive actions require confirmation ("Type CONFIRM to proceed")
  • Failures are reported clearly with next steps

GitHub's Hubot, Slack's bot framework, and similar tools make this straightforward. The pattern works for any infrastructure operation that benefits from visibility.

Idempotency

Automated runbooks should be safe to re-run. The most painful runbook failures are the ones where step 3 of 10 errored, and you cannot tell if running it again from the top is safe.

Patterns for idempotent runbooks:

  • Check-and-act. Every step checks the state before acting. If already done, skip.
  • Resumable. The script can pick up from a checkpoint if it failed mid-way.
  • Reversible. Each step has a corresponding "undo" step in case of failure.
# Idempotent example
if aws rds describe-db-instances --db-instance-identifier "$NEW_PRIMARY" \
   --query 'DBInstances[0].DBInstanceClass' --output text \
   | grep -q "^db.r6g"; then
  echo "Already upgraded, skipping"
else
  aws rds modify-db-instance ...
fi

The script can be re-run after a failure without making things worse.

Testing Runbooks

A runbook tested only in production is a recipe for incident escalation.

Two testing patterns:

Staging exercise. Run the runbook in staging on a schedule. Catches drift between expected and actual environment.

Game day execution. Run the runbook during a game day exercise (covered in another post). Tests the runbook against a real failure.

Untested runbooks are aspirations, not infrastructure. They will not work the first time they are needed.

What Belongs in a Runbook

Useful runbook content:

  • The trigger (when to run this)
  • Prerequisites (what state must be true)
  • The command to run
  • Expected output and how to interpret it
  • What to do if it fails
  • Who to escalate to if you cannot resolve

Less useful:

  • Long prose introductions
  • History of why the runbook exists
  • Detailed system architecture
  • Hypothetical scenarios

The runbook is read under pressure. Cut everything that does not help the responder right now.

What Doesn't Belong in a Runbook

Things that look like runbook material but should live elsewhere:

  • Architecture explanations. Belongs in design docs.
  • Capacity planning. Belongs in dashboards.
  • Incident postmortems. Belongs in the incident archive.
  • Vendor support contacts. Belongs in a dedicated escalations page.

When everything ends up in the runbook, finding the actual steps becomes the problem.

Versioning and Discoverability

Runbooks in a repo benefit from:

  • A clear naming convention (runbooks/database/promote_replica.sh)
  • An index file listing them with their triggers
  • Linking from alert configuration to the relevant runbook
  • A simple search (a grep across the repo is often enough)

The alert that wakes the engineer should link directly to the runbook. "Database connection saturation" alert links to runbooks/database/scale_connections.sh (which has the wiki description of when and why at the top).

The Honest Limit

Not every operation can or should be automated. Some legitimate reasons to keep a runbook manual:

  • The operation is rare enough that automation cost exceeds savings
  • The operation requires judgment that resists encoding
  • The operation has variable steps depending on context

For these, a clearly-written wiki runbook is fine. Do not force everything into automation; the discipline is to automate the things that benefit from it.

A Practical Adoption Path

For a team adopting runbook automation:

  1. Pick the top 5 most-run runbooks. These have the highest automation payoff.
  2. Express them as scripts in a repository. Plain bash is fine; do not over-engineer.
  3. Document the script's purpose and usage at the top.
  4. Have someone other than the author execute the script in staging. Catches assumptions.
  5. Link the alerts to the runbook scripts.

Week one investment, ongoing payoff. The next on-call shift is the test.

The Outcome

Teams that automate runbooks well notice three changes:

  • On-call shifts become less stressful (the responder has tools, not just instructions)
  • Incidents resolve faster (consistent execution, no manual mistakes)
  • Onboarding new on-call engineers is easier (the script teaches the procedure)

The investment pays back within months. The runbook automation is the artifact; the operational confidence is the actual value.


Looking at runbooks that have lived in a wiki for years and are due for upgrade? We help teams move from wiki-based procedures to scripts that survive the 2 AM test. scopeforged.com

Share this article

Related Articles

Need help with your project?

Let's discuss how we can help you build reliable software.