The section on fire department metrics does an incredible job of explaining why MTTR isn’t a useful metric.
Brent Chapman
Full disclosure on this one, the author is my former (awesome) boss, and I think I may be that staff engineer they mentioned…
Reid Savage — Honeycomb
Rootly grapples with how to evolve their PR review process with the increase in LLM-generated code.
The harder bugs hide in context, shared boundaries, and rollout paths.
Quentin Rousseau — Rootly
The ladder does not end at SEV-3. It does not end at SEV-4. It ends somewhere below, in a category we have decided not to name.
Tim Irving
The shift to distributed systems over monoliths can increase complexity. We don’t need to go back to the old way, says this article, but we do need to build with operability in mind.
Vivek Kadam — Communications of the ACM
Every service catalog is declared by hand and drifts from reality within weeks. That was a productivity tax when humans read it. Now that agents act on it, a stale catalog is a production risk.
I’m not so sure about the solution offered, but the risk is real. Humans can already act on incorrect service catalog information, but agents can do it much more quickly.
Spiros Economakis
What you’ll learn in this post isn’t a success story, it’s a learning journey. We’ll walk through the architecture decisions that enabled scale, the production challenges that tested those decisions, the optimization methodology that guided us through, and the lessons that apply to any distributed system.
Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva, and Nathan Fisher
A summary of Uptime Labs’s Incident Fest, with a lot of interesting tidbits around AI, incidents, and the evolving interplay between them.
Sam Salter — Uptime Labs
