SRE Weekly Issue #527

A message from our sponsor, Planetscale:

Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

Explore PlanetScale

Amazing idea: negative time to detection. It’s when you know an incident is coming even before the impact actually begins. Some incident management and metrics systems aren’t designed to track it, and some incident processes ignore or even penalize it.

  Tim Irving

The key question is, when that new code breaks in production at 3am, how well can the on-call engineers debug it?

  Brent Chapman

This article focuses heavily on how Honeycomb built and tested a plan for what I can personally assure you must have been a very complex migration. I especially like how they drew lessons from their incident late last year.

  Josh Parsons — Honeycomb

An in-depth introduction to sharding in PostgreSQL. Sure, you probably know all about sharding, but I definitely learned some interesting bits from this one even so.

  Ben Dicken — PlanetScale

  This article is published by this issue’s sponsor, but their sponsorship did not influence its inclusion in the newsletter.

I started out ready to hate this one, but by the end, I came around; there’s a lot to think about. It’s about how coding agents change the economics of saying no or yes to taking on projects.

Cheap to write is not the same as cheap to own

  Dalia Abuadas — GitHub

In this commercial airliner near miss, a system designed for reliability (alpha floor protection) was the cause of an incident (uncommanded pitch down), giving an excellent case study in automation.

Do we know when it works, how it works, how to get the most out of it, and how to find a way around it if it turns against us?

  Mentour Pilot

Conventional wisdom says a database makes a bad queue and you should reach for a real broker but these folks went the other direction, tearing out RabbitMQ in favor of polling MongoDB. They included a section at the end of everything they had to build themselves to make polling work.

  Pavel Zavialov — DevOps.com

A detailed description of their approach, including how they tested it and the weak points they’re still iterating on.

  Akhilesh Rao Meesala — HackerNoon

Updated: July 26, 2026 — 9:15 pm
A production of Tinker Tinker Tinker, LLC Frontier Theme