SRE Weekly Issue #528

A message from our sponsor, Planetscale:

Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

Explore PlanetScale

Spotify has had some difficulty around podcast publishing, and they shared this analysis of the worst incident.

  Jim Whitehead, Ulrik Mikaelsson, John Lagomarsino, and Saunak Jai Chakrabarti — Spotify

…and here’s where it gets interesting. This post shares the user point of view on the Spotify issues, including fact-checking their published timeline.

  Gergely Orosz — The Pragmatic Engineer

New incident role unlocked: the incident tech lead. I enjoyed the description of the interplay between the tech lead and the incident commander.

  Brent Chapman

This one goes hard: if you try to reduce your incident count, your system will become less reliable, not more. Aim for more incidents, handled well.

  Tim Irving

There’s some brutal honesty in here that I find refreshing, especially around the impact on incidents and incident response.

  Liz Fong-Jones — Honeycomb

Here’s what I’ve learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

  Karan Nagarajowda — Uptime Labs

agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it.

   Ajay Devineni — DZone

The premise: read replicas can help you scale read load, but they introduce complexity. The article goes into the problems they ran into and how they dealt with them.

  Johanna Larsson — incident.io

Updated: August 2, 2026 — 9:56 pm
A production of Tinker Tinker Tinker, LLC Frontier Theme