The Problem
An alert fired. No one had to answer it, so before long, no one did.
For this platform, that was the real failure. Not that a service broke, but that nothing was watching on a schedule, and nothing guaranteed the right person ever heard about it.
What We Understood First
01
Why Failures Went Unnoticed
No schedule checking health. No defined path from failure to person.
02
A Watchdog, Every 30 Seconds
Every registered service polled continuously. An incident opens the moment one goes quiet.
03
A Ladder, Not a Broadcast
Each tier is a role and a wait time. Unanswered, it climbs. Acknowledged, it stops.
04
Every Channel, Independently
Email, SMS, WhatsApp, push, tried on their own, not just one and done.
05
Configurable, Not Fixed
Every rule's ladder, roles, and wait times are set by the team, not hardcoded.
“A page that always means something. That's the whole product.”
Enterprise-scale platform, escalation configured per service
What We Didn't Automate
The system decides what to watch, when to open an incident, when to climb a tier. All mechanical.
People define the rules themselves, what counts as a trigger, which role owns it, how long each tier waits, and a human must acknowledge and resolve every incident.
Tensions Worth Naming
Alert fatigue.
Alerts that never resolve train people to stop looking at them. This one only escalates while it stays unacknowledged, spread by round-robin so no single person carries the load.
Trust at the top.
A page that reaches leadership routinely stops meaning anything. Here, the top tier is reached only after every lower tier fails to acknowledge, so a page that reaches leadership is, by construction, a confirmed incident nobody has picked up yet.
If You're Facing This Too
Check first whether your alerts actually require a person to acknowledge them. An alert nobody has to close is an alert nobody trusts, and it teaches people to ignore the rest.
Make escalation climb by role when nobody acknowledges, and save the top tier for exactly those cases, so a page there always means act.
What's Next
More of the platform's services under the same watchdog, and incidents tied to their own history for context.
None of this needs a bigger on-call roster as coverage grows, the same ladder scales to more services without more people watching dashboards. And most monitoring tools stop at the alert; guaranteeing an escalation path with defined roles and wait times is what actually closes the loop.
IN PERSPECTIVE
This platform needed something sturdier than more alerts: confidence that every one of them would reach somebody, and stay theirs until it was handled.
If you're navigating layered challenges and
want a thinking partner — let's think together.
START A CONVERSATIONContinue Exploring
DRAWING INTELLIGENCE
BOM Automation
How one panel-building firm stopped every BOM depending on one engineer's eye.
Read →MARGIN INTELLIGENCE
Pricing Intelligence
How one industrial distributor stopped pricing by memory across hundreds of local companies.
Read →NEGOTIATION INTELLIGENCE
Procurement Intelligence
How one distribution network stopped conceding on price with nothing asked in return.
Read →