SRE Weekly Issue #530

A message from our sponsor, Planetscale:

Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.

→ Get started with PlanetScale for just $5/mo

We may improve velocity by handing off tasks to LLM agents, but can that impact resilience?

  Courtney Nash — Resilience in Software Foundation

Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too.

  Brent Chapman

A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.

  Zak van der Merw

A harrowing incident story underlining the importance of expertise and experience.

  Hamed Silatani — Uptime Labs

An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.

  Bill Duncan

Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.

Bonus: they include links to several write-ups of related incidents.

  TokenTimer

Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.

Whoa, cool trick!

  Pragya Mehta and Sai Samant — Stripe

Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.

There’s an interactive simulation of their algorithm midway through that’s fun to play with!

  Patrick Hamann and Mike Fisher — incident.io

Updated: August 16, 2026 — 11:44 pm
A production of Tinker Tinker Tinker, LLC Frontier Theme