Designing an On-Call Rotation That People Don't Hate

Philip Rehberger Oct 2, 2026 7 min read

Build rotations that respect humans: handoffs, follow-the-sun, and compensation models.

On-call rotations get designed once, then live with for years. A poorly-designed rotation produces burnout, turnover, and "we cannot find anyone to take call." A well-designed rotation produces engineers who feel responsible for their systems without resenting it.

This post is the patterns that hold up — what to optimize for, what tradeoffs to accept, and how the rotation itself becomes part of how the team thinks about their work.

What On-Call Is Really For

The conventional answer is "responding to production incidents." Useful framing, but incomplete. On-call is also:

  • The mechanism by which engineers feel the consequences of their architectural decisions
  • A forcing function for alert and observability quality
  • A real-time learning opportunity about how the system actually behaves
  • A career-development experience for understanding production at scale

A rotation designed only for "respond to alerts" misses the second-order benefits. A rotation designed for all of them produces stronger engineers and better systems.

The Anatomy of a Rotation

Six dimensions define how an on-call rotation works:

Frequency. How often a given engineer is on call. Weekly? Every three weeks? Once a quarter?

Duration of shift. A week? 24 hours? 12-hour blocks?

Scope. What systems and services does the on-call cover? One team's services? Multiple teams?

Tier. Single-tier (everyone responds to everything), two-tier (primary plus escalation), multi-tier (L1 / L2 / L3)?

Compensation. Paid extra for on-call? Time off in lieu? Just part of the job?

Backup. A secondary on-call? A team chat? A specific person to call?

Each dimension has real tradeoffs. The right answers depend on incident volume, team size, system complexity, and what the team can tolerate.

Healthy Rotation Patterns

A rotation that works long-term tends to have:

  • Weekly shifts rather than 24-hour rotations. Single-night shifts are easier on individuals; weekly rotations create ownership.
  • Frequency around every 4–6 weeks. Less than every 4 weeks burns people out; more than 6 weeks means losing familiarity.
  • Two-tier setup. Primary responds; secondary backs up the primary if unreachable. The secondary's existence reduces the primary's anxiety.
  • Clear scope. "You are on-call for these specific services, not for everything." Limits blast radius and cognitive load.
  • Compensation. Some recognition that on-call is work. Money, time off, or both.

A rotation with these properties is sustainable. A rotation with a weekly shift, no secondary, no compensation, and ambiguous scope produces turnover.

The Burnout Risks

Three patterns reliably produce burnout:

Too-frequent shifts. Every other week is too often. Engineers never feel "off." Sleep gets disrupted regularly. Family life suffers.

Too-loud alerts. When the pager fires three times per shift, every shift is exhausting. Worse, the team stops trusting alerts.

Lone responder. No secondary, no team chat. When something happens, one person carries the entire weight.

If any of these are true in your rotation, fix them before fixing anything else.

Alert Hygiene

The biggest determinant of how painful on-call is is not the rotation design — it is the alert volume.

A page in the middle of the night should mean something. Most teams accumulate alerts over time that no longer mean anything:

  • Symptoms of conditions that recovered themselves
  • Alerts for components that have been retired
  • Thresholds set during a different era of load
  • Alerts that fire only when someone is awake to fix them

A useful rotation health metric: how many pages per shift? Three is too many. One per shift is reasonable. Zero is great but not the goal — some pages are part of the job.

If pages per shift trends up, the rotation will degrade no matter how good the schedule is. Investment in alert quality is investment in on-call sustainability.

Compensation Models

Common patterns:

  • No extra pay, "it's part of the job." Works only for relatively quiet rotations with mature teams. Fails when expected pages are common.
  • Flat stipend per shift. $200–$500 per week of on-call. Simple, predictable, fair.
  • Per-page compensation. Hours of pay per alert. Discourages alert flap by aligning incentives.
  • Time in lieu. A day off after a particularly bad shift. Often combined with another mechanism.

For US-based teams: a flat stipend is most common and most appreciated. For teams that get paged often: per-page or comp-time-after-bad-shifts is more sustainable.

The compensation matters less than the principle: on-call is recognized as work, not just expected.

Follow-the-Sun

For globally distributed teams, follow-the-sun rotations are appealing — handoff at the end of one region's day to the next region's morning. No one carries night shifts.

The catch: it requires real coverage in three regions, with enough engineers in each to run a rotation. Most teams do not have this. A "follow-the-sun" with one engineer in Asia is a single point of failure pretending to be a rotation.

When it works, follow-the-sun is genuinely better than night shifts. When it does not, it is "one tired engineer in a weird timezone."

Documentation Requirements

On-call works when the responder can act. That means:

  • Runbooks for every alert. What does this alert mean, what to check, what to do.
  • Service ownership maps. Who owns what. Where to escalate if it is not "your" service.
  • Access to the systems needed. Read access to production, the ability to run runbooks, dashboards bookmarked.
  • Out-of-band contacts. Vendor support phone numbers, escalation paths, on-call for adjacent teams.

A rotation without these is asking engineers to improvise under pressure. The first time someone gets paged for a service they have never seen, the cost is paid.

Handoff Practices

Shift transitions are when context gets lost. The handoff at the end of a week should include:

  • Open incidents that have not fully resolved
  • Tickets created during the shift
  • Anomalies noticed that have not been investigated
  • Upcoming deploys or migrations during the next shift

Async written handoffs work better than sync calls. A markdown note in a shared place, read at the start of the next shift, beats a 15-minute call that nobody remembers.

Senior/Junior Mix

A pure-junior rotation is dangerous; a pure-senior rotation is unsustainable.

Practical patterns:

  • Pair junior with senior for the first several shifts. The junior is primary; the senior is secondary.
  • Specific time before solo. "Three pairing shifts before you go solo" gives juniors the practice without the full pressure.
  • Senior escalation. A clear path from primary to a senior engineer for anything outside the runbook.

This is also part of how on-call develops engineers. A junior who has been on-call for a year understands production in ways that classroom training cannot teach.

Vacation and Coverage

Engineers take vacation. The rotation has to support that without manual scrambling.

The cleanest pattern: pre-schedule swaps. When an engineer schedules vacation, they coordinate a swap with someone before they leave. The rotation never has gaps.

Avoid the "asking in the team chat the day before" pattern. It works but adds anxiety and disrupts the schedule.

Tools That Help

The major tools:

  • PagerDuty — most common, expensive but polished.
  • Opsgenie (Atlassian) — competitive features, often cheaper.
  • Better Stack, Squadcast, Spike — newer alternatives with smaller footprints.
  • Slack-only — for very small teams. Slack as the entire pager system works for low-volume teams.

The tool matters less than the rotation discipline. A poorly-designed rotation in PagerDuty is just as miserable as one in Slack.

When to Restructure

Signals that the current rotation is failing:

  • Engineers actively trying to avoid on-call slots
  • High turnover correlated with on-call burden
  • Frequent failures to find a responder
  • The phrase "I had a really bad on-call week"
  • Senior engineers carrying disproportionate load

When these appear, the rotation is not the problem on its own — it is a symptom of alert volume, scope, or staffing. Address the underlying cause; the rotation will follow.

The Real Goal

A healthy on-call rotation is one where engineers can do it, talk about it like any other work, and not associate it with dread. The system is not "we never get paged" — that is unrealistic. The system is "when we get paged, the response is competent, the workload is reasonable, and the team improves the system afterwards."

Get there and on-call becomes a feature of engineering culture, not a tax on it.


Designing or revising an on-call rotation that has worn the team down? We help teams structure rotations and alert hygiene that sustains over years, not quarters. scopeforged.com

Share this article

Related Articles

Need help with your project?

Let's discuss how we can help you build reliable software.