I’ve racked up more on-call rotations than I care to count. Some of them hollowed me out. Some bored me into questioning my life choices. And a tiny handful—the ones worth remembering—made me trust my teammates enough to actually put my phone on silent and sleep like a normal person. But here’s what sticks: your on-call setup isn’t a process problem. It’s a mirror. It shows you exactly what your team gives a damn about. Not the fluffy stuff on the careers page. Not the CTO’s quarterly platitudes. The raw, unfiltered, 3 a.m. truth about your engineering culture.
Want to size up a team fast? Skip the code reviews. Ignore the sprint boards. Watch the pager. Who takes the hit? What kind of noise trips the wire? And what actually happens after someone squelches the alarm? That’s where the real values hide, usually under a pile of good intentions and ignored runbooks.

The Pager Is a Lie Detector
Every company swears they worship at the altar of reliability. Dashboards glow with “uptime.” Job ads brag about five nines like a badge of honor. But on-call is where the slogan meets the grind, and the gap can be a mile wide. I’ve landed on teams where the pager only screams for customer-facing disasters—a database blip gets a shrug until Tuesday because “it’s just internal.” That isn’t reliability. That’s putting on a show for the people who sign the checks.
Real reliability means caring about the failures that don’t make the CEO’s phone buzz. When your on-call engineer is told to mute staging alerts without a second thought, your team is announcing that shipping fast matters more than building something solid. When they’re expected to fix the mess alone, without dragging anyone else out of bed, you’re betting on heroics over collaboration. The pager doesn’t spin the story. It just beeps. You draw the conclusions.
What Silent Alerts Say About Your Priorities
Silent alerts are the passive-aggressive sticky notes of on-call. Somebody, somewhere, decided this particular dumpster fire wasn’t worth a human’s eyeballs, so it gets logged and forgotten. I joined a team once that had over 200 silent alerts firing every single day. The ops channel looked like a morgue—just bot messages scrolling into the abyss. When I asked what the deal was, the lead shrugged and said, “We’d never sleep otherwise.” Translation: we’d rather protect our REM cycles than fix the garbage that threatens them.
Look, I’m not anti-sleep. Sleep is holy. But if you’re silencing alerts because the noise is unbearable, you don’t have a monitoring problem. You have a values problem. You’ve decided that the long-term health of the system—and the sanity of whoever’s on call next month—takes a back seat to tonight’s uninterrupted dreams. That’s a choice. Just be honest about it.

Rotation Design Is a Moral Document
The way you structure a rotation is a map of who holds the power. A team that spreads the pain evenly—seniors, juniors, backend grumps, frontend wizards—values shared ownership. A team that keeps a dedicated ops crew on a permanent, soul-crushing rotation values specialization. And frankly, a caste system. The second approach isn’t automatically wrong; sometimes you need a SWAT team for the nuclear stuff. But let’s not slap a “we’re all in this together” sticker on it.
I’ve watched startup founders personally field every single page. That’s not dedication. That’s a control complex with a hoodie. The team learns that the founder’s ego matters more than anyone else’s growth. On the other end of the circus, I’ve seen the new hire tossed into primary on-call during week two. That’s not onboarding. That’s hazing with a monitoring dashboard. The value there? “Sink or swim, kid.” Both setups are broken, just painted different colors of dysfunction.
The Follow-the-Sun Lie
Follow-the-sun sounds so civilized: spread on-call across time zones so nobody loses a full night. I’ve watched it get twisted into a cover for deeper rot. When the handoff between regions is a joke—no runbooks, zero context, just a Slack message grunting “your turn”—the team values geography over clarity. The engineer in Mumbai inherits a steaming pile of unresolved alerts at 9 a.m., while the engineer in New York logs off feeling virtuous. That’s not a rotation. That’s a relay race where the baton is actively on fire.
Doing follow-the-sun right takes real investment: overlapping hours, shared dashboards, and a culture where dumping an active incident on the next region feels like a small betrayal. If your team won’t do that work, you’re just outsourcing the misery to whatever time zone is cheapest.
Incident Response Reveals the Blame Game
Post-incident reviews are where the masks slip off entirely. A team that runs blameless postmortems—and I mean actually blameless, not the fake kind where you say the magic word but still highlight who shipped the bad commit—values learning. A team that kicks off with “whose fault was this?” values covering your own ass. I’ve been in rooms where the first question after an outage was, “Did we miss an alert threshold?” That’s a team that believes in systems thinking. I’ve also been in rooms where the first question was, “Why didn’t Dave test this?” That’s a team that believes in finding a warm body to take the hit.
The worst flavor I ever tasted: a team that threw a party for the hero who fixed the incident and never once asked why the thing broke in the first place. The hero got a shoutout in the weekly newsletter. The root cause—a deployment pipeline made of toothpicks and hope—survived another six months untouched. That team valued hero stories more than boring reliability work. So they got more heroes. And way more incidents.
Escalation Paths as Power Maps
Stare at your escalation policy for a minute. Who gets the call when the primary can’t fix it? If it’s always the same two grizzled senior engineers, your team values institutional knowledge stuffed inside human skulls over actual documentation. Those two people are walking single points of failure, and the team is weirdly okay with that because writing stuff down is harder than pinging Dave again. I’ve been that engineer. It feels important right up until you realize you’re just a human wiki with a pager tan and a permanent twitch.
If the escalation path is just “yell into the ops channel and hope someone’s awake,” your team values chaos. That can almost work in a five-person startup where everyone lives in the codebase. It falls apart spectacularly at twenty people, when the junior frontend dev gets paged for a database corruption because they were the only green dot in Slack. Chaos doesn’t scale. Neither does dumb luck.

Compensation Is the Loudest Signal
Nothing screams “we value this” quite like cash. If your on-call engineers get paid extra—and I mean real money, not a cold pizza at 2 a.m.—the team values their time. If on-call is just “part of the gig” with zero extra compensation, the team values saving a buck over preventing burnout. I’ve seen companies offer “comp time” that nobody can ever use because the sprint deadlines are carved in stone. That’s not compensation. That’s a guilt coupon you’ll never redeem.
One crew I worked with gave on-call engineers a bonus for every incident-free week. Sounds brilliant, right? Until you notice it incentivizes people to quietly close incidents without digging for the root cause. The mean time to resolve dropped, sure. But the same bugs kept boomeranging back. The team valued a clean report over a clean system. The bonus was basically hush money.
The Tools You Pay For
Do you shell out for PagerDuty, or do you route alerts through a free Slack integration that falls over every other month? Do you invest in monitoring that surfaces real signal, or do you rely on a homegrown script that someone’s former intern cobbled together in 2019? The tools you buy—or refuse to buy—reflect the value you place on the on-call experience. A team that won’t spend $50 a month on a decent alerting platform but drops thousands on a team outing is investing in morale theater, not actual morale.
What Healthy On-Call Actually Looks Like
It’s not complicated, but finding it in the wild is like spotting a unicorn. A healthy on-call rotation has a few non-negotiables: the pager fires rarely, and when it does, the alert is something a human can act on. The rotation includes everyone who ships code, and nobody is chained to the pager for more than a week at a time. Incidents kick off blameless reviews that produce concrete fixes, not a round of awkward apologies. And the team gets compensated—in cash or real, usable time off—for the hours they’re tethered to the damn thing.
That setup values reliability, fairness, learning, and respect for human time. It’s not rocket science. It’s just hard because most teams would rather not face the trade-offs. They want the benefits of on-call without the costs. They want reliability without ever slowing down the feature factory. They want happy engineers without paying for their lost sleep. The pager will tell you exactly which side they’ve chosen. You just have to listen.
FAQ
What’s the biggest red flag in an on-call rotation?
A pager that fires nonstop for stuff that isn’t urgent. That tells me the team can’t tell the difference between noise and signal, which means they don’t value systematic reliability work. It’s a cultural failure wearing a technical hat.
Should junior engineers be on call?
Yes, but never alone. Pair them with a senior engineer for the first few rotations. If you throw them into the deep end solo, you’re not building skills—you’re building a deep reservoir of resentment. The team that does this values a “trial by fire” fantasy over actual mentorship.
How do I convince my team to pay for on-call?
Stop framing it like you’re asking for a favor. On-call is work—it strangles your freedom, trashes your sleep, and piles on stress. If the company won’t compensate for that, they’re announcing exactly how much they value your personal time. Walk them through the math on burnout and turnover. If they still balk, polish your résumé.
Is a “no on-call” culture possible?
For some products, sure—if you can genuinely tolerate downtime without customers howling. But if your service matters to someone at 3 a.m., a human has to be available. The question isn’t whether you have on-call; it’s whether you design it to be humane. A team that brags about “no on-call” but then panics when something breaks on a Saturday values denial over honesty. And that’s a rough way to run a shop.