The Demon of Oncall [8dde0150]
Software engineers contend with oncalls in different ways. Some find companies to peace out where there is no oncall load, some pull the Google and try to build follow-the-sun rotations, some move to teams where the oncall burden is light. I've chosen the worst of all these options: choosing to run into the fire.
This isn't about how to run an oncall, that's largely banal crisis response.
In many ways this is sort of my own fault, WhatsApp has a culture of reiability, but also, like many companies, uneven and uncaring oncalls. So thus the holy crusader must rise, for there are two types of projects: projects that are green and new, and projects that are dragons that have slain careers. My project is shaping up to be slaying the oncall culture dragon.
So what the fuck is even my oncall? At WhatsApp, I'm primarily oncall for Core Messaging (your messages), Media (your cat photos and videos), and some other bits and bobs. This means that my job is roughly equivilant to the reiable transport of light.
But what are the actual demons? There's actually only one demon: the demon of acedia. Because the vast majority of oncall expectations are "whatever is previously there", there is strong tendency to respect authoritarial intent. Imagine a situation where an alarm fires, but nothing notable actually occurs, and the oncall marks it as transient.
Why was this alarm threshold set at 50%? Nobody knows! But now we don't want to bump it too far away, to 75%, because we'd really want to go to 55, then 65, then 75, because smaller steps are an easier change, even those the alert may not be providing any value at all. But no, we can't remove it either, because what if this alert is meant to catch something? Surely we can't argue with a potential outage in the future! Ignore that by the time we bumped it to 75, this alert would be useless, but all the while this sucks away at our ability to actually rethink the alert.
It might've been better off to just remove the alert, and see which other part squeals next time. Many of the alerts that are created in the wake of an outage are leftover scabs, at a certain point, we need to have them fall off.
But nobody wants to be responsible for an outage. Or be perceived as being responsible for the outage. But this is a fundalmental risk one must take when attempting to drive oncall change.
> talk about how we need to create, and rethink a lot of it > software relaibility is about building empathy with the machine so you can feel how it breaks, not sticking test lines in every single line into a sick patient > build faster tools and intutition, not stick alarms in every single place possible
> discursive effort on oncall > Do you truly believe that poverty erases sin, that communal property and the exaltation of sex and the sensuality of the dance and the rejection of all authority and the unrestrained life of vagabonds in the forest, on the beaches, and along the highways could supplant and even overcome the established order?