SRE Weekly Issue #537

A message from our sponsor, WarpBuild:

WarpBuild runs your GitHub Actions jobs on runners that are 2x faster at half the cost. Change one line in your workflow file and keep everything else. Get $50 in free credits to try it now

Try WarpBuild free

When an Anthropic engineer says they’re seeing problems when using Claude for incident response, it’s well worth a read. The story seems to be improving, but it’s worth reading both this talk summary and The Register’s summary of a previous talk.

  Sylvain Kalache, summarizing talks by Alex Palcuie (Anthropic)

Another excellent story of shaving bytes at scale, with a fun coda: how do you safely change the hashing scheme in your massively distributed cache at scale without invalidating the whole thing?

  Kevin Guthrie, Mariia Iurchenko, Zaidoon Abd Al Hadi, and Ivan Babrou — Cloudflare

I can’t even count how many times I’ve pulled up the Wikipedia article with the nines chart. Maybe I’ll bookmark this interactive tool instead.

  fivenines

I’m impressed by the guardrails and safeguards they built into this system, and the automated fallback.

  Violetta Pidvolotska and Naveen Mareddy — Netflix

As the first in a series, this article sets the stage, going into why we autoscale and the autoscaling dimensions available in Kubernetes. I like the emphasis on the balance between cost and reliability.

  drmorr — Applied Computing Research Labs

Third-party SDKs speed up development, but they also introduce performance, security, reliability, and maintenance risks that teams must actively manage.

   Satyam Nikhra — DZone

This one’s worth paying attention to if you might need to scale up any time soon.

  Gergely Orosz

Here’s Lorin’s take on an incident I linked here a couple weeks back. I love this way of framing it:

The report attributes the incident to two errors:

  1. correctly following an incorrect procedure
  2. incorrectly following a correct procedure

  Lorin Hochstein

SRE Weekly Issue #536

A message from our sponsor, Planetscale:

PlanetScale is headed to SREcon26 in Dublin this October. Swing by our booth to grab some merch, catch a live demo, and chat with the team behind the world’s fastest databases. We can’t wait to see you there.

→ If you’re not at SREcon but still want to learn how PlanetScale can reliably scale your databases, get in touch.

Common threads seen across incident write-ups from many companies.

  Lorin Hochstein

How we built a production architecture handling 1,900+ messages per second and tens of thousands of concurrent connections with Kafka, Redis Pub/Sub, .NET Channels, and SSE.

  Mohsen Rajabi — ITNEXT

This is a thorough tour of the design decisions that separate a background job system that works in development from one that survives production.

  Shreshtha Saha — Techgenyz

This question kicked off a great comment section:

Has anyone gotten an automated RCA setup to actually nail root cause without a person doing the final synthesis, or is that still mostly aspirational marketing from vendors?

  u/Acrobatic_Refuse8100 and many others — reddit

Do you have a way to slow down traffic to your database during an incident? This one has a great explanation of why you need one.

  Lorin Hochstein

Through a fictitious case study, this article shows how to go about building a reliable service with an agentic component.

FYI the last ~quarter or so is a sales pitch, but the preceding majority of the article isn’t.

  Arpio

Rollback sounds great in theory, but it doesn’t always work. The CircleCI example really hits hard.

This reminds me of a classic article from the now-defunct company Skyliner, You Can’t Have a Rollback Button.

  Balu Kambala

Yes, this is a walkthrough of a vendor’s product, but the architecture is genuinely interesting and it reads like an engineering explainer rather than a sales pitch. I especially liked the wrinkle where one shard has to be designated authoritative so that custom type OIDs stay consistent across the whole cluster.

  Harshit Gangal — PlanetScale

  This article is published by this issue’s sponsor, but their sponsorship did not influence its inclusion in the newsletter.

SRE Weekly Issue #535

A message from our sponsor, Planetscale:

Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale:

  • Online schema changes
  • Version upgrades with no downtime
  • Better connection pooling
  • Resharding
  • Graceful planned and unplanned failovers

→ Neki is available today. Learn more.

The early-career perspective here is so valuable! I love the way they explain everything, and the extension on their project is really nifty.

  Anthony Oparaocha — incident.io

I enjoyed the evaluation of AI agents in the frame of control theory and Ashby’s law. When your control system is incredibly complex, it can be hard to reason about what it does.

  Lorin Hochstein

Their scale is such that every byte shaved off each cache entry saves 250 GB fleet-wide.

  Sebastiaan Neuteboom — Cloudflare

Here’s a good one to dig into from GCP. It’s especially interesting to see what they knew when they wrote their preliminary report vs their final report.

  Google

There’s a really cool idea in this article and the original post it refers to: overlaying histograms to show how a latency distribution has changed over time. I’m still wrapping my head around the kernel density graphs, though.

  Bill Duncan

A sudden scale-up put pressure on CoreDNS which didn’t have autoscaling configured. After reading this incident report, you may also want to check out Lorin Hochstein’s analysis.

  Buildkite

What can the shape and progression of impact during an incident tell us about how our systems are working?

  Bala Subrahmanyam Kambala — WeAreDevelopers

Monitoring isn’t enough, you need to make sure the renewed TLS certificate actually gets deployed, and a lot can get in the way. I like this concept of “CertOps”.

  TokenTimer — HackerNoon

SRE Weekly Issue #534

A message from our sponsor, Planetscale:

Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale:

  • Online schema changes
  • Version upgrades with no downtime
  • Better connection pooling
  • Resharding
  • Graceful planned and unplanned failovers

→ Neki is available today. Learn more.

The better these tools become at resolving routine incidents, the less practice human responders will get. And when an ambiguous, high-severity incident comes in that automation cannot solve, responding engineers will be in trouble.

I love the concept of comprehension debt described in this article.

  Sylvain Kalache

Another excellent write-up by my favorite writer of incident write-up write-ups. I especially like that last section.

Lorin also wrote more the next day.

  Lorin Hochstein

A VP’s question […] lands with the weight of the org chart behind it, and everyone in the channel feels it.

I especially liked the sections toward the end on how to mitigate a leader’s impact.

  Brent Chapman

Lowe’s has an intensive (and intense) SRE practice that provides client teams with a framework of tools to ensure reliability.

  Raghuprasanth Ravichandran, Joe Praveena A, and Pavan Palagiri

How far will the impact of an incorrect automated decision spread? Will it get integrated into the system and influence later decisions?

  Sai Sandeep Koneti — HackerNoon

A great primer on how pod disruption budgets work and why they’re critical for reliability.

  Andre Newman — Gremlin

They took a systematic approach using test-driven development. I enjoyed not only learning what went well, but where they ran into trouble.

  Edvinas Janusevicius — Checkly

Great advice if you’re training early-career engineers on incident response. This advice reminds me of some of the techniques that folks used with me when I was starting out.

  Karan Nagarajagowda — Uptime Labs

SRE Weekly Issue #533

A message from our sponsor, Planetscale:

Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

→ Explore PlanetScale

What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company’s main web page for a sudden uptick in traffic.

  Brent Chapman

What an interesting incident! I recommend reading Azure’s write-up before reading Lorin’s excellent analysis.

  Lorin Hochstein

Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.

   Varsha Ganesh — DZone

I love this concept of a “political incident”:

The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.

And ouch, I felt this bit:

You have spent forty minutes of the incident on the severity field.

  Tim Irving

Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.

  Sai Joshitha Kathari — HackerNoon

I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you’re trying to roll out a fix during an incident.

  Mike Thompson and Daniel Esponda — Datadog

Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn’t a simple drop-in replacement.

  Samuel Yeboah, Francesco Di Chiara and Mingliang Liu — Netflix

Can we replace human code review with LLM-based reviews? This article lays out what an LLM can’t replicate, and I’d argue that these are the pieces that matter most for reliability.

  John Allspaw — Adaptive Capacity Labs