SRE Weekly Issue #535

A message from our sponsor, Planetscale:

Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale:

  • Online schema changes
  • Version upgrades with no downtime
  • Better connection pooling
  • Resharding
  • Graceful planned and unplanned failovers

→ Neki is available today. Learn more.

The early-career perspective here is so valuable! I love the way they explain everything, and the extension on their project is really nifty.

  Anthony Oparaocha — incident.io

I enjoyed the evaluation of AI agents in the frame of control theory and Ashby’s law. When your control system is incredibly complex, it can be hard to reason about what it does.

  Lorin Hochstein

Their scale is such that every byte shaved off each cache entry saves 250 GB fleet-wide.

  Sebastiaan Neuteboom — Cloudflare

Here’s a good one to dig into from GCP. It’s especially interesting to see what they knew when they wrote their preliminary report vs their final report.

  Google

There’s a really cool idea in this article and the original post it refers to: overlaying histograms to show how a latency distribution has changed over time. I’m still wrapping my head around the kernel density graphs, though.

  Bill Duncan

A sudden scale-up put pressure on CoreDNS which didn’t have autoscaling configured. After reading this incident report, you may also want to check out Lorin Hochstein’s analysis.

  Buildkite

What can the shape and progression of impact during an incident tell us about how our systems are working?

  Bala Subrahmanyam Kambala — WeAreDevelopers

Monitoring isn’t enough, you need to make sure the renewed TLS certificate actually gets deployed, and a lot can get in the way. I like this concept of “CertOps”.

  TokenTimer — HackerNoon