Amazing idea: negative time to detection. It’s when you know an incident is coming even before the impact actually begins. Some incident management and metrics systems aren’t designed to track it, and some incident processes ignore or even penalize it.
Tim Irving
The key question is, when that new code breaks in production at 3am, how well can the on-call engineers debug it?
Brent Chapman
This article focuses heavily on how Honeycomb built and tested a plan for what I can personally assure you must have been a very complex migration. I especially like how they drew lessons from their incident late last year.
Josh Parsons — Honeycomb
An in-depth introduction to sharding in PostgreSQL. Sure, you probably know all about sharding, but I definitely learned some interesting bits from this one even so.
Ben Dicken — PlanetScale
This article is published by this issue’s sponsor, but their sponsorship did not influence its inclusion in the newsletter.
I started out ready to hate this one, but by the end, I came around; there’s a lot to think about. It’s about how coding agents change the economics of saying no or yes to taking on projects.
Cheap to write is not the same as cheap to own
Dalia Abuadas — GitHub
In this commercial airliner near miss, a system designed for reliability (alpha floor protection) was the cause of an incident (uncommanded pitch down), giving an excellent case study in automation.
Do we know when it works, how it works, how to get the most out of it, and how to find a way around it if it turns against us?
Mentour Pilot
Conventional wisdom says a database makes a bad queue and you should reach for a real broker but these folks went the other direction, tearing out RabbitMQ in favor of polling MongoDB. They included a section at the end of everything they had to build themselves to make polling work.
Pavel Zavialov — DevOps.com
A detailed description of their approach, including how they tested it and the weak points they’re still iterating on.
Akhilesh Rao Meesala — HackerNoon
