SRE Weekly Issue #531

A message from our sponsor, Planetscale:

PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.

→ See the benchmarks

Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall.

  Brent Chapman

What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems?

  Fred Hebert

Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces.

  Ashwini Dave — DZone

If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say?

It’s all about people. I really enjoyed the quote from Charles Billings on principles for automation.

  Steven Shorrock

Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE.

  Bill Duncan

Some big names in this Q&A, and they share a couple of delicious morsels.

  Sam Salter — Uptime Labs, with John Allspaw and Beth Adele Long

A super-engaging deep-dive.

This investigation is a useful reminder: running boring technology in a non-standard way is a risk.

  Alex Chan — Tailscale

A handy guide on topology constraints in Kubernetes, with a worked example.

  Andre Newman — Gremlin