Common threads seen across incident write-ups from many companies.
Lorin Hochstein
How we built a production architecture handling 1,900+ messages per second and tens of thousands of concurrent connections with Kafka, Redis Pub/Sub, .NET Channels, and SSE.
Mohsen Rajabi — ITNEXT
This is a thorough tour of the design decisions that separate a background job system that works in development from one that survives production.
Shreshtha Saha — Techgenyz
This question kicked off a great comment section:
Has anyone gotten an automated RCA setup to actually nail root cause without a person doing the final synthesis, or is that still mostly aspirational marketing from vendors?
u/Acrobatic_Refuse8100 and many others — reddit
Do you have a way to slow down traffic to your database during an incident? This one has a great explanation of why you need one.
Lorin Hochstein
Through a fictitious case study, this article shows how to go about building a reliable service with an agentic component.
FYI the last ~quarter or so is a sales pitch, but the preceding majority of the article isn’t.
Arpio
Rollback sounds great in theory, but it doesn’t always work. The CircleCI example really hits hard.
This reminds me of a classic article from the now-defunct company Skyliner, You Can’t Have a Rollback Button.
Balu Kambala
Yes, this is a walkthrough of a vendor’s product, but the architecture is genuinely interesting and it reads like an engineering explainer rather than a sales pitch. I especially liked the wrinkle where one shard has to be designated authoritative so that custom type OIDs stay consistent across the whole cluster.
Harshit Gangal — PlanetScale
This article is published by this issue’s sponsor, but their sponsorship did not influence its inclusion in the newsletter.
