The early-career perspective here is so valuable! I love the way they explain everything, and the extension on their project is really nifty.
Anthony Oparaocha — incident.io
I enjoyed the evaluation of AI agents in the frame of control theory and Ashby’s law. When your control system is incredibly complex, it can be hard to reason about what it does.
Lorin Hochstein
Their scale is such that every byte shaved off each cache entry saves 250 GB fleet-wide.
Sebastiaan Neuteboom — Cloudflare
Here’s a good one to dig into from GCP. It’s especially interesting to see what they knew when they wrote their preliminary report vs their final report.
There’s a really cool idea in this article and the original post it refers to: overlaying histograms to show how a latency distribution has changed over time. I’m still wrapping my head around the kernel density graphs, though.
Bill Duncan
A sudden scale-up put pressure on CoreDNS which didn’t have autoscaling configured. After reading this incident report, you may also want to check out Lorin Hochstein’s analysis.
Buildkite
What can the shape and progression of impact during an incident tell us about how our systems are working?
Bala Subrahmanyam Kambala — WeAreDevelopers
Monitoring isn’t enough, you need to make sure the renewed TLS certificate actually gets deployed, and a lot can get in the way. I like this concept of “CertOps”.
TokenTimer — HackerNoon
