SRE Weekly Issue #537

A message from our sponsor, WarpBuild:

WarpBuild runs your GitHub Actions jobs on runners that are 2x faster at half the cost. Change one line in your workflow file and keep everything else. Get $50 in free credits to try it now

Try WarpBuild free

When an Anthropic engineer says they’re seeing problems when using Claude for incident response, it’s well worth a read. The story seems to be improving, but it’s worth reading both this talk summary and The Register’s summary of a previous talk.

  Sylvain Kalache, summarizing talks by Alex Palcuie (Anthropic)

Another excellent story of shaving bytes at scale, with a fun coda: how do you safely change the hashing scheme in your massively distributed cache at scale without invalidating the whole thing?

  Kevin Guthrie, Mariia Iurchenko, Zaidoon Abd Al Hadi, and Ivan Babrou — Cloudflare

I can’t even count how many times I’ve pulled up the Wikipedia article with the nines chart. Maybe I’ll bookmark this interactive tool instead.

  fivenines

I’m impressed by the guardrails and safeguards they built into this system, and the automated fallback.

  Violetta Pidvolotska and Naveen Mareddy — Netflix

As the first in a series, this article sets the stage, going into why we autoscale and the autoscaling dimensions available in Kubernetes. I like the emphasis on the balance between cost and reliability.

  drmorr — Applied Computing Research Labs

Third-party SDKs speed up development, but they also introduce performance, security, reliability, and maintenance risks that teams must actively manage.

   Satyam Nikhra — DZone

This one’s worth paying attention to if you might need to scale up any time soon.

  Gergely Orosz

Here’s Lorin’s take on an incident I linked here a couple weeks back. I love this way of framing it:

The report attributes the incident to two errors:

  1. correctly following an incorrect procedure
  2. incorrectly following a correct procedure

  Lorin Hochstein