
Nobody signs up for an engineering gig dreaming about the on-call rotation. We dream about building things that scale, shipping features people actually want, and writing code so clean it hurts. Then someone hands you a pager, points you at a Slack channel called #incidents, and a runbook last updated when the lead architect still had a full head of hair. Just like that, you’re not an engineer anymore. You’re a night-shift paramedic for a codebase that won’t stop hemorrhaging.
I’ve survived more on-call rotations than I’ve had bad hangovers, and I’ve arrived at one solid, ugly truth: your on-call setup isn’t a process. It’s a mirror. It reflects what your team, your manager, and your whole organization actually value. Not the glossy stuff on the careers page, but the things they’ll drag you out of bed at 3 a.m. for. Everything else is a sales pitch.
Want to know if your company gives a damn about reliability, work-life balance, or basic human dignity? Skip the employee handbook. Look at the on-call schedule. Look at who gets paged, how often, and what happens when they finally crawl back under the covers. The rot lives in the rotation.
The 3 a.m. Test: What Wakes You Up?
Let’s start with the most obvious signal: the alerting criteria. Every team says they care about uptime. Every manager frowns earnestly during incident reviews and mumbles about “prevention.” But what actually triggers a page? That’s where the bullshit unravels.
I was once on a team where the on-call engineer got paged every time a particular microservice’s latency ticked past 200ms. Not 500ms. Not a full second. Two hundred milliseconds. The service wasn’t even customer-facing; it was an internal caching layer that could chug along at twice that speed without a single user blinking. But the architect had skimmed a blog post about “performance budgets” and decided 200ms was the hill he’d die on. So, night after night, some sorry soul got shaken awake because a garbage collection hiccup added 20ms to a request nobody was waiting for.
What did that team value? Not reliability. Not user experience. They valued the architect’s ego. They valued sounding clever in meetings. The tab got picked up by engineers who learned to dread their jobs.
Another team I worked with ran a critical payment service. Their threshold? “Page if the error rate climbs above 5% for ten minutes.” Five percent! A one-in-twenty failure rate on transactions that moved actual money, and they’d let it stew for ten whole minutes before waking someone. By then, the finance department was already sharpening their pitchforks. What did that team value? Avoiding on-call discomfort. They’d rather hemorrhage revenue than interrupt an engineer’s sleep. Sounds almost kind until you realize it just postponed the agony until the 9 a.m. meeting where everybody screamed at each other.
The thresholds you set are a blunt statement of priorities. Tight thresholds scream, “We value catching problems early, even if we burn out our people.” Loose thresholds mutter, “We value uninterrupted sleep, even if it means bigger fires later.” Neither choice is automatically wrong. But if your team hasn’t had a straight-up conversation about which one you’re picking, you’re just drifting into a culture by accident. And accidental cultures are almost always a slow poison.

The Runbook Litmus Test: Documentation as a Cultural Artifact
Runbooks. In theory, a runbook is a step-by-step guide for diagnosing and fixing common nightmares. A gift from your well-rested daytime self to your bleary 3 a.m. self. In practice, it’s a confession.
I’ve thumbed through runbooks that stretched to 40 pages, obsessively maintained, complete with flowcharts and decision trees. I’ve also seen runbooks that consisted of a single line: “Restart the service and if that doesn’t work, call Dave.” Both tell you exactly what the team values.
The 40-page runbook team valued resilience and knowledge sharing. They grasped that the person on call might be new, exhausted, or simply not the expert on that particular service. They put in the time so that anyone could handle an incident. That’s a team that values its people. End of story.
The “call Dave” team valued hero culture. They probably thought they valued speed and pragmatism. What they actually valued was having a single wizard whose skull contained all the tribal knowledge. Dave gets to feel indispensable. Dave also gets to burn out, rage-quit, or get flattened by a bus, leaving the team utterly screwed. They didn’t value resilience; they valued the illusion of it, propped up by one overworked human.
And here’s the sharp bit: the state of your runbooks is never just about documentation habits. It’s about whether your team views on-call as a shared duty or a punishment. If you don’t update runbooks after incidents—if you don’t treat that as actual work deserving time and respect—you’re telling your engineers their suffering isn’t worth preventing. You’re saying, “We’d rather you figure it out from scratch next time than spend an hour writing down what you learned.” That’s not a technical failure. That’s a values failure.
Compensation and the Unspoken Contract
Money talks, and nowhere does it yell louder than in on-call compensation. How your company handles this tells you whether they see on-call as part of the gig or as a sacrifice that demands acknowledgment.
Some places fork over extra cash for on-call shifts. Some offer time off in lieu. Some do neither and pretend it’s just “baked into the salary.” I’ve collected paychecks from all three, and the difference is night and day—often literally.
When you get paid extra for on-call, the company is drawing a clear line: this is above and beyond your normal duties. We see the burden. We’re not going to pretend that being unable to leave town without a laptop is just another Tuesday. That acknowledgment changes the emotional math. You still despise the pager, but at least you feel seen.
When there’s no extra compensation, the company is making an equally clear statement: your time isn’t yours. We own you. The salary covers whatever chaos we decide to lob at you, including 2 a.m. alerts about a server someone forgot to patch. This is the express lane to resentment, and resentment is the silent team-killer. People don’t quit over on-call pay. They quit because they realize the company views them as a resource to drain, not a human to respect.
Then there’s the murky middle: time off in lieu. This can work if it’s actually honored. But I’ve watched too many teams where “take the morning off after a rough call” morphs into “well, standup is at 9, but maybe you can sneak out early.” That’s not compensation. That’s a bad joke. If your team values work-life balance, the compensation will be concrete and non-negotiable. If it’s fuzzy and left to a manager’s whim, the value is only on paper.

Who Gets to Be On Call?
This is the question that makes managers squirm. Glance at your on-call roster. Is it the same three battered souls every month? Are new hires tossed into the deep end during their first week? Is there a shadowy cabal of senior engineers whose names never darken the schedule?
The makeup of the on-call rotation exposes the real power structure of your team. If only junior engineers carry the pager, the message is blunt: on-call is grunt work. It’s beneath the senior folks. They’ve paid their dues, and now they get to sleep. This creates a two-tier system where the least experienced people handle the most stressful part of the job. It’s not just unfair; it’s reckless. When a genuine disaster strikes, you want your most battle-scarred people on the front line, not the new grad still googling how to SSH into a box.
If senior engineers are on call but rarely get paged because they’ve actually built systems that don’t fall over, that’s a different beast. That’s a sign of maturity. But that’s not the typical story. Usually, the seniors have quietly arranged to be “too critical” for on-call, and the weight lands on the people least able to push back.
And what about managers? If your engineering manager isn’t on the rotation—or at least in an escalation path that genuinely gets used—then your team values hierarchy over solidarity. I’m not saying every manager needs to lug a pager, but if they’ve never tasted that 3 a.m. dread, they will never truly understand what they’re asking of their team. The best managers I’ve worked with pulled occasional on-call shifts, not because they were ace troubleshooters, but because they wanted to stay connected to the pain. That’s a value statement right there.
The Blamelessness Mirage
Every modern engineering team loves to trumpet their “blameless culture.” It’s one of those phrases that’s been repeated so often it’s become hollow, like “we’re agile” or “we take security seriously.” Your on-call process will show you whether it’s real.
Here’s the test: what happens after an incident? If the postmortem is a calm, curious autopsy of what went wrong and how to prevent it, congratulations—you might actually have a blameless culture. But if the postmortem is a thinly disguised interrogation where someone’s hunting for “who screwed up,” you don’t have blamelessness. You have a blame culture with a better PR team.
I’ve sat in postmortems where the lead engineer asked, “Why did you push that config change without testing?” and the room turned to ice. That’s not a question. That’s an accusation. The real question should be, “Why did our system allow a config change to be pushed without testing?” The first question assumes the problem is a person. The second assumes the problem is the system. Which one your team reaches for first tells you everything about whether they actually value learning or just value looking blameless on their LinkedIn profiles.
On-call exposes this because incidents are inherently stressful. When the adrenaline is hammering and the CEO is pacing the war room, the mask slips. You glimpse what people really believe. If your team can stay curious and supportive in the guts of a Sev1, you’ve built something rare. If they start eating each other, you’ve built a pressure cooker that will eventually blow.
Alert Fatigue and the Slow Death of Giving a Damn
There’s a beast called alert fatigue, and it’s exactly what it sounds like. When you get paged too often for nonsense, you stop caring about the pages. You silence your phone. You ignore Slack. Then one day, you miss a real alert, and the whole system craters for hours.
Alert fatigue isn’t a technical problem. It’s a cultural one. It breeds when a team values “being informed” over “being effective.” Someone wires up alerts for every conceivable metric, and nobody has the spine to prune them because they’re terrified of missing something. So they miss everything instead.
The cure is brutal prioritization, and that demands a team that values focus and clarity over covering their ass. You need someone with the authority—and the nerve—to say, “This alert is noise. We’re killing it. If something bad happens because of that, I’ll own the fallout.” How many teams have that person? Almost none. Because most teams value dodging blame more than doing solid work. So the alerts pile up, the engineers burn out, and everyone stands around wondering why the on-call rotation is a revolving door.
Rotations as a Forcing Function
Here’s the thing I’ve come to believe after years marinating in this nonsense: on-call isn’t just a necessary evil. It’s a forcing function for engineering culture. It forces you to stare at the gap between what you say and what you do. It forces you to decide whether you’re going to invest in reliability or just yap about it. It forces you to treat your teammates like actual humans or watch them walk out the door.
If your on-call rotation is a disaster, your team is a disaster. You can have all the slick architecture diagrams and pristine code you want, but if the human system for keeping the lights on is broken, the technical system will eventually follow. People aren’t interchangeable cogs. They’re the ones who write the code, swat the bugs, and answer the pages. If you treat them like resources, they’ll act like it—and resources don’t care if the site goes dark.
So study your rotation. Look at who’s on it, what yanks them awake, and what happens afterward. If you don’t like what you see, don’t start by fiddling with alert thresholds or polishing runbooks. Start by asking your team what they actually value. Then ask whether the on-call process mirrors that. The answer will be uncomfortable, but it’s the only conversation that counts.
Frequently Asked Questions
How do we know if our on-call rotation is fair?
Fairness isn’t about identical hours clocked; it’s about the weight carried. Check if some people get paged far more often than others during their shifts. Check if certain services are “cursed” and always land on the same unlucky engineers. Also, ask your team anonymously: does on-call feel like a shared load or a punishment? If the answer leans toward punishment, your rotation isn’t fair, no matter how pretty the schedule looks on a spreadsheet.
Should we compensate on-call with money, time off, or nothing?
Something always beats nothing. Cash sends a clear signal that the company values your time. Time off can work if it’s genuinely protected—meaning you’re not expected to sneak a peek at email or join meetings during your recovery window. “Nothing” is only acceptable if on-call is genuinely rare and uneventful, which almost never happens. If you’re getting paged more than once a month on average, compensation should be on the table. If your company balks, they’re telling you what they value, and it’s not you.
What’s the single biggest mistake teams make with on-call?
Treating it like a purely technical puzzle. Teams obsess over monitoring tools, alert thresholds, and runbook templates, but they ignore the human stuff: burnout, resentment, and the fear of being blamed. The biggest mistake is assuming that a good on-call process springs from good tooling. It doesn’t. It springs from a culture that genuinely cares about reliability and the people responsible for it. Everything else is just buying a fancier pager.
How do we reduce alert fatigue without risking major outages?
Start by sorting every alert into three buckets: “page right now,” “can wait until morning,” and “belongs on a dashboard, not a pager.” Be vicious. If you’re uncertain, default to less paging—you can always escalate later. Then, every quarter, review the alerts that actually fired and ask: did this page lead to a meaningful action? If not, demote it. This review has to be a team ritual, not a side chore for the one person who actually cares about monitoring.