SRE Weekly Issue #529

A message from our sponsor, Planetscale:

Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.

→ Get started with PlanetScale for just $5/mo

It’s not enough to define an incident process. You have to spin up and maintain an entire incident management program.

  Brent Chapman

Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans’. It’s especially interesting that increasing agent usage has not correlated with decreasing human usage.

  Austin Parker — Honeycomb

I like the approach here, especially measuring both the positive and negative outcomes.

Good SRE practice is about evidence, not enthusiasm.

   Neel Shah — DZone

The kernel’s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.

  Ray Chen — Railway

…we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.

  TW Lim — Antithesis

I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.

  Omar Ghader

First time I’ve heard of systemd’s PrivateTmp feature. Neat!

  Chris Siebenmann

I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.

It’s short (just a table), but it definitely made me think.

  Lorin Hochstein

Updated: August 9, 2026 — 10:28 pm
A production of Tinker Tinker Tinker, LLC Frontier Theme