SRE Weekly Issue #532

A message from our sponsor, Planetscale:

PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.

→ See the benchmarks

Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn’t a good idea.

  Brent Chapman

Ethics are relative, right? This article is full of genuinely useful tips and framings.

  Thomas A. Limoncelli — ACM Queue

Whoa. It’s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist.

  Jorrick Sleijster — Adyen

For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.

  Sridhar Rajarao

Traditional observability monitors execution. LLM observability must monitor behavior.

  Barnadeep Bhowmik

This one has a lot of great detail on how their approaches to quota management failed and how they iterated.

  Dhyanam Vaidya, Prathamesh Deshpande, and Mike Ma — Uber

This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation.

  Robert Barron