SRE Weekly Issue #534

A message from our sponsor, Planetscale:

Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale:

  • Online schema changes
  • Version upgrades with no downtime
  • Better connection pooling
  • Resharding
  • Graceful planned and unplanned failovers

→ Neki is available today. Learn more.

The better these tools become at resolving routine incidents, the less practice human responders will get. And when an ambiguous, high-severity incident comes in that automation cannot solve, responding engineers will be in trouble.

I love the concept of comprehension debt described in this article.

  Sylvain Kalache

Another excellent write-up by my favorite writer of incident write-up write-ups. I especially like that last section.

Lorin also wrote more the next day.

  Lorin Hochstein

A VP’s question […] lands with the weight of the org chart behind it, and everyone in the channel feels it.

I especially liked the sections toward the end on how to mitigate a leader’s impact.

  Brent Chapman

Lowe’s has an intensive (and intense) SRE practice that provides client teams with a framework of tools to ensure reliability.

  Raghuprasanth Ravichandran, Joe Praveena A, and Pavan Palagiri

How far will the impact of an incorrect automated decision spread? Will it get integrated into the system and influence later decisions?

  Sai Sandeep Koneti — HackerNoon

A great primer on how pod disruption budgets work and why they’re critical for reliability.

  Andre Newman — Gremlin

They took a systematic approach using test-driven development. I enjoyed not only learning what went well, but where they ran into trouble.

  Edvinas Janusevicius — Checkly

Great advice if you’re training early-career engineers on incident response. This advice reminds me of some of the techniques that folks used with me when I was starting out.

  Karan Nagarajagowda — Uptime Labs