We may improve velocity by handing off tasks to LLM agents, but can that impact resilience?
Courtney Nash — Resilience in Software Foundation
Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too.
Brent Chapman
A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.
Zak van der Merw
A harrowing incident story underlining the importance of expertise and experience.
Hamed Silatani — Uptime Labs
An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.
Bill Duncan
Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.
Bonus: they include links to several write-ups of related incidents.
TokenTimer
Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.
Whoa, cool trick!
Pragya Mehta and Sai Samant — Stripe
Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.
There’s an interactive simulation of their algorithm midway through that’s fun to play with!
Patrick Hamann and Mike Fisher — incident.io
