The better these tools become at resolving routine incidents, the less practice human responders will get. And when an ambiguous, high-severity incident comes in that automation cannot solve, responding engineers will be in trouble.
I love the concept of comprehension debt described in this article.
Sylvain Kalache
Another excellent write-up by my favorite writer of incident write-up write-ups. I especially like that last section.
Lorin also wrote more the next day.
Lorin Hochstein
A VP’s question […] lands with the weight of the org chart behind it, and everyone in the channel feels it.
I especially liked the sections toward the end on how to mitigate a leader’s impact.
Brent Chapman
Lowe’s has an intensive (and intense) SRE practice that provides client teams with a framework of tools to ensure reliability.
Raghuprasanth Ravichandran, Joe Praveena A, and Pavan Palagiri
How far will the impact of an incorrect automated decision spread? Will it get integrated into the system and influence later decisions?
Sai Sandeep Koneti — HackerNoon
A great primer on how pod disruption budgets work and why they’re critical for reliability.
Andre Newman — Gremlin
They took a systematic approach using test-driven development. I enjoyed not only learning what went well, but where they ran into trouble.
Edvinas Janusevicius — Checkly
Great advice if you’re training early-career engineers on incident response. This advice reminds me of some of the techniques that folks used with me when I was starting out.
Karan Nagarajagowda — Uptime Labs
