Welcome to this week’s Uptime Sync. This issue covers how a broken DNSSEC rollover took down Albania’s .AL.AL.AL domain, how Modal scales to 111 million concurrent sandboxes in seconds, and the lessons Netflix learned while building service topology at scale. We also dig into modern storage I/O bottlenecks, why job queues are deceptively difficult to get right, and how an unpatched Argo CD vulnerability can lead to Kubernetes cluster takeover through cache poisoning.

On the database and reliability front: why PostgreSQL teams use strict memory overcommit to avoid the OOM killer, how to make PostgreSQL prune partitions even when queries filter on non-partitioned columns, and the operational details that matter when systems are under real pressure.

On the tutorial front: running self-hosted LLMs on Kubernetes with vLLM, using Copilot to profile and fix Java heap sizing, integrating OpenTelemetry Collectors with Kubernetes and VictoriaMetrics, building a VPN with Tailscale, shipping blue/green deployments with Argo Rollouts, and building container networking from scratch.

Newsworthy Reads

Tutorials of the Week

Projects of the Week

Join 1,000+ engineers staying ahead of the curve

Every week, Uptime Sync brings you:

  • Outage postmortems from Netflix, Cloudflare, Pinterest & more

  • Hands-on DevOps & SRE tutorials

  • Production-ready tools & open-source projects

👋 Find me on Twitter | Linkedin | Connect 1:1

Thank you for supporting this newsletter. Consider sharing this post with your friends.

Y’all are the best.

Keep Reading