When Maya Chen joined Harbr as VP of Engineering, the on-call rotation was quietly destroying the team. Two senior engineers had already left. More were close. Here’s how she fixed a problem that turned out to be much less about alerts, and much more about what an engineering culture actually values.
The first thing Maya Chen did when she joined Harbr was ask to see the on-call logs. Not the incident reports, not the postmortems, the raw logs. Every page, every alert, every 3am notification that had fired in the previous six months. What she found was not a reliability problem. It was a tax, a slow, compounding levy being charged against the attention and sleep and goodwill of every engineer on the rotation.
Harbr builds data infrastructure tooling for mid-market companies: the kind of product where uptime genuinely matters and where a real incident at the wrong moment can cause real downstream pain for customers. The on-call rotation existed for good reasons. But the alerts had metastasized.
Over eighteen months of rapid growth, the monitoring system had accumulated hundreds of rules, many of them added in the heat of a specific incident and never revisited. Engineers were being woken up for memory usage crossing a threshold that had been set conservatively in 2021 and never recalibrated. For database replication lag that was, in practice, always self-correcting. For CPU spikes that correlated perfectly with a scheduled batch job that ran every night at 2am and had never once caused a customer-facing issue.
“The alerts were technically correct,” Chen says. “That was almost the whole problem. Nobody could argue that any individual alert was wrong to fire. But collectively, they were telling the team that everything was urgent, which meant nothing was.”
Two senior engineers had left in the four months before Chen arrived. Exit interviews, read carefully, pointed to on-call as a significant factor in both departures. A third, one of the team’s most respected infrastructure engineers, had quietly told his manager he was interviewing elsewhere. Chen had six months before a Series B fundraising process in which investors would be asking pointed questions about team stability.
The Wrong Diagnosis
The conventional response to a noisy on-call rotation is a tooling conversation. Better alerting platforms. Smarter routing. Automated runbooks. Chen had seen that playbook applied before and understood its limitations. Better tools for managing bad alerts are still bad alerts. The underlying problem wasn’t that the team lacked the infrastructure to handle pages more efficiently. It was that the organisation had never forced itself to answer a harder question: what is actually worth waking someone up for?
“On-call dysfunction is almost always a prioritisation problem wearing a technical costume,” she says. “You can re-route alerts more cleverly, but if you haven’t decided what matters, you’re just shuffling the same noise around.”
Chen’s first move was to convene what she called an alert audit, not a technical review, but a values exercise. She gathered the six engineers on the current rotation and asked them to go through every alert that had fired in the previous ninety days and answer a single question for each one: if this alert fires at 3am and you don’t respond until 8am, what is the actual customer impact?
The exercise took two sessions of two hours each. The results were stark. Of the 340 distinct alert types that had fired in the audit period, 71 percent either had no measurable customer impact if left unaddressed overnight, or would resolve themselves before a human could meaningfully intervene. Twelve percent were genuinely urgent. The remaining 17 percent sat in a grey zone that required further discussion.
Nobody in the room was surprised by the findings. That, Chen noted, was the most revealing thing about the exercise. “Everyone already knew which alerts were noise,” she says. “The organisation just hadn’t created a legitimate process for acting on that knowledge.”
Making the Implicit Explicit
What followed the audit was, in some ways, the harder part. Identifying noise is relatively straightforward. Getting an organisation to act on that identification requires navigating a particular kind of institutional anxiety: the fear that downgrading an alert is the same as saying the thing it monitors doesn’t matter.
Chen addressed this directly. She introduced a three-tier classification system for all alerts: page-worthy (genuine customer impact requiring immediate human response), ticket-worthy (something that needs attention in the next business day), and log-worthy (captured for trend analysis but requiring no immediate action). Every alert in the system had to be assigned to one of the three tiers, with written justification and a named owner responsible for reviewing the classification quarterly.
The classification process was deliberately slow and deliberate. Teams reviewed their own alerts, proposed tiers, and then defended those proposals in a thirty-minute cross-team review. The cross-team element mattered: it created accountability, prevented individual teams from quietly inflating the importance of their own systems, and, perhaps most valuably, forced conversations between teams that rarely talked about how their systems affected each other.
“Some of our most useful conversations happened in those reviews,” says Priya Nair, one of Harbr’s senior backend engineers and an early sceptic of the process. “We realised that one team was paging on-call for a dependency that another team had actually made more resilient three months earlier. Nobody had updated the alert. It had just kept firing.”
What Changed
The reclassification process took eight weeks from the first audit session to full implementation. In that time, the team reviewed 340 alert types, reclassified 241 of them as ticket-worthy or log-worthy, and eliminated 34 alerts entirely as genuinely obsolete. The number of alerts capable of generating an out-of-hours page dropped by 60 percent.
The impact was measurable almost immediately. In the quarter before the audit, engineers on the rotation were averaging 4.2 out-of-hours pages per week. In the quarter after, that number fell to 1.6. Mean time to acknowledge dropped, because the pages that did fire were ones the team recognised as genuinely important and responded to with urgency rather than resignation.
The engineer who had been quietly interviewing elsewhere stayed. In a conversation with Chen three months after the reclassification, he described the change in direct terms: “It’s not that the job got easier. It’s that it started feeling like the organisation respected my time.” That distinction, between ease and respect, is the one that matters most for retention. Engineers leave when they feel that their attention is being treated as a cheap resource. The audit was, at its core, an act of institutional acknowledgement that it wasn’t.
In the two quarters following implementation, Harbr saw zero on-call-related attrition. The classification framework has since been adopted by three other teams inside the company, including the data platform and security teams, neither of which was part of the original process.
The Deeper Principle
Chen is careful not to oversell the framework. “This is not a permanent fix,” she says. “Alerts will drift back toward noise if you don’t maintain discipline. The quarterly reviews aren’t optional.” The classification system is a living document, not a one-time cleanup, and she treats it as such.
But the deeper principle she draws from the experience is one that applies well beyond on-call rotations. Fast-growing engineering organisations accumulate technical debt of all kinds, in codebases, in architectures, in processes. Alert debt is a specific and particularly damaging form of that accumulation, because it is paid not in future maintenance cost but in present human wellbeing. Every spurious page is a small withdrawal from the trust and energy of the person who receives it.
“Engineers are not infinitely elastic,” Chen says. “You can ask them to absorb a lot. But there’s a version of high standards that tips into extractiveness, and once people feel that line has been crossed, you don’t get their trust back easily.”
The Series B closed at a valuation that reflected, among other things, a stable and high-performing engineering team. The investors who asked about team stability got the answer they were looking for. The engineer who had been planning to leave is now leading Harbr’s infrastructure reliability programme. And the on-call logs, reviewed quarterly, remain the first thing Chen looks at when she wants to understand how the team is really doing.
Note: This is a spec writing sample. Harbr, Maya Chen, and all characters and companies depicted are fictional.

Comments
Post a Comment