Every Tuesday - Deep dives, architecture lessons, and real engineering stories.
Every Saturday - The best DevOps, SRE, Cloud, AI, tools, tutorials, and projects from the week.
📬 Missed This Week’s Uptime Sync?
This week’s best DevOps, SRE, Cloud, Kubernetes, database, AI infrastructure, and production engineering reads:
How Cloudflare saved another 100 TB of RAM with math and Rust
How Linux readahead works—and when to change it
Why a Postgres migration passed every test but failed with real data
Why every AI agent may need its own computer
How n8n was deployed, maintained, and migrated on EKS
PostgreSQL row-level security performance in practice
How an OpenAI DNS jailbreak actually worked
Plus: debugging stuck Kubernetes rollouts, understanding environment variables, using Claude Code from anywhere, scaling Docker on Railway, managing node drains with Pod Disruption Budgets, and tools for SQLite, AI security, OpenTofu, dev containers, and QUIC.
I’m thinking of creating practical DevOps/SRE interview guides for Kubernetes, Cloud, Linux, SRE, DevOps, and more.
I know, I know.
Debugging is a boring topic to read.
I could have written about whatever is trending this week - JEV, Gemini 4, some new AI model, agent framework, or another shiny tool everyone is talking about.
That probably would've gotten more clicks.
But if you're trying to become a genuinely good engineer, I'd rather you understand this.
Because once you get good at debugging, the real world becomes much easier.
Tools will change. Frameworks will change. The stack you use today might be completely different three years from now.
But production will still break.
An API will suddenly start returning 401.
A deployment that worked yesterday will start throwing 500.
A container will be running perfectly fine but somehow have no logs.
A script will fail with an error that makes absolutely no sense at first.
And in those moments, knowing another Kubernetes command or memorising another system design diagram won't save you.
How you think will.
Start with what you actually know
Imagine someone tells you:
The app is broken.
That doesn't give you much to work with.
Instead, write down what is actually happening.
Request: GET /profile
Expected: 200 OK with profile JSON
Actual: 401 Unauthorized
Scope: happens in my local test client only
Started: after I changed the client environment variables
Recent change: renamed API_TOKEN to AUTH_TOKENNow you have something you can investigate.
This distinction matters.
A 401 Unauthorized is something you know happened.
The authentication config is broken.
That's only a guess.
Other possible reasons could be:
the
Authorizationheader was never sentthe token expired
you're calling the wrong endpoint
the server is reading a different environment variable
a proxy changed the request
Debugging gets much easier when you stop treating your first guess like the answer.
For a small script, the same idea applies.
Instead of:
My Python script doesn't work.
Write:
Command: python totals.py
Expected: total is 15
Actual: TypeError: unsupported operand type(s)
Input: quantity came from a text fileYou now have a specific failure that you can reproduce.
That's already a much better starting point.
When something breaks, do this
You don't need a complicated debugging framework.
Use these five steps.
1. What exactly is broken?
Be specific.
Bad:
The API is broken.Better:
GET /profile should return 200 but returns 401.Bad:
Docker logs aren't working.Better:
The container is running, the request fails, but docker logs returns nothing.The more clearly you describe the problem, the easier it becomes to decide what to check next.
2. What changed?
A lot of bugs appear shortly after something changed.
Check things like:
a new deployment
configuration changes
environment variables
dependency upgrades
different input
DNS changes
certificates
permissions
network rules
infrastructure changes
This doesn't mean the latest change is automatically the cause.
It just gives you somewhere useful to start looking.
3. Pick one possible reason
Don't try to explain the entire problem at once.
Pick one thing that could explain what you're seeing.
For example:
Maybe the Authorization header isn't being sent.That's enough.
4. Find one way to check it
Now ask:
If this is actually the problem, what should I see?
If the client isn't sending an Authorization header, inspecting the request should show that the header is missing.
So run a check that answers exactly that question.
5. Let the result choose your next step
This is where good debugging starts to feel very different from guessing.
If the header is missing:
> investigate the client.
If the header exists:
> stop touching the client and move to the server.
Each result should make your search area smaller. That's the goal.
Suppose your local API should return:
200 OKbut you're getting:
401 UnauthorizedA random approach might look like this:
edit authentication code
restart the server
change an environment variable
disable an auth check
restart everything
try again
Maybe it starts working. But which change fixed it? You don't know.
And worse, maybe disabling the auth check was the thing that made it work.
That's not really a fix.
Instead, start with one possible reason:
Maybe the client isn't sending the Authorization header.Now ask:
If that's true, what should I see?You should see that the outgoing request doesn't contain an Authorization header.
So inspect the request:
curl -v http://localhost:8080/profileWith curl -v, outgoing request lines start with > and response lines start with <.
Look for something like:
> GET /profile HTTP/1.1
> Host: localhost:8080
> Authorization: Bearer ...If the header isn't there, you just narrowed the problem down.
Now check the client.
Maybe your code expects:
API_TOKENbut your environment only contains:
AUTH_TOKENThat's a useful finding.
You went from:
Authentication is broken.to:
The client isn't sending the token because it's reading the wrong environment variable.Much better.
But what if the header is present? Good.
Your original guess was wrong. Stop changing client code. Move to the server.
Check:
server authentication logs
token validation
token expiry
environment variables
middleware
proxy configuration
A wrong guess is still useful when it tells you where not to look.
One thing to remember: curl -v shows what the client tried to send. If there is a proxy between the client and server, the request could still be changed later.
Also be careful before sharing verbose output. It can contain tokens, cookies, and other secrets.
Example 2: docker logs is empty
Here's another common one.
Your container is running.
Requests to the application are failing.
So you run:
docker logs my-appAnd get nothing useful back.
It's easy to conclude:
The application didn't log any errors.
But that's not necessarily true.
A better question is:
Where is this application actually writing its logs?
docker logs normally shows output Docker captured from the container process, usually stdout and stderr.
But an application might instead write logs to:
/var/log/app/error.logor send them somewhere else entirely.
So start with one possible reason:
Maybe the application writes errors to a file instead of stdout or stderr.Now inspect what Docker actually started:
docker inspect my-app --format '{{.Path}} {{json .Args}}'If the image has a shell, you can inspect the container:
docker exec -it my-app shThen check the application's configuration.
Look for:
log file paths
logging configuration
process arguments
environment variables
external logging agents
Maybe you discover:
log_file=/var/log/app/error.logNow you know why docker logs looked empty.
The application was logging. Docker just wasn't showing those logs.
The exact commands will differ between applications and container images.
Some containers don't even include a shell. That's fine.
The useful part is the order:
> docker logs is empty
> Where does the app write logs?
> What process is actually running?
> What logging configuration is it using?
> Where should I look next?That's debugging.
Don't confuse fixing the symptom with finding the cause
This becomes especially important in production.
Imagine you deploy a new release.
A few minutes later:
POST /checkoutstarts returning:
500 Internal Server ErrorCustomers are affected. You roll the release back. The error rate drops. Users can check out again.
But what did you actually learn?
You know:
Rollback changed the system and the errors stopped.You do not automatically know:
The new application code caused the problem.The rollback may also have changed:
traffic patterns
container instances
cache state
connections
configuration
dependency behaviour
timing
So keep two separate questions in your head.
Question 1: How do we stop users from being affected?
That might mean:
rollback
disable a feature
route traffic elsewhere
scale something up
fail over
restart unhealthy instances
Question 2: Why did this happen?
That investigation can continue after the system is stable.
During an incident, write down what you actually know:
Impact:
Checkout requests return 500 for some users
Start time:
14:07 UTC
Scope:
One region
Mitigation:
Rolled back release 2025.03.18.2
Observed result:
Error rate returned to normal
Root cause:
UnknownNotice the last line.
Root cause: UnknownThere's nothing wrong with that.
It's much better than pretending your first theory is already confirmed. This is why incident response and root-cause analysis are different jobs.
During an outage, your first priority may be getting the service healthy again.
Understanding exactly why it happened can come next.
A command should answer a question
One habit I see beginners develop is collecting debugging commands.
top
free -m
df -h
netstat
docker logs
kubectl describe
journalctlThese are useful tools.
But knowing commands isn't the same thing as knowing how to debug.
Before running one, ask:
What am I trying to learn from this?
Instead of:
Let me run top.Think:
The service became slow.
Could CPU saturation explain it?
I'll check CPU usage.Instead of:
Let me run df -h.Think:
The application can't write new files.
Could the disk be full?
I'll check filesystem usage.Instead of:
Let me look at Kubernetes events.Think:
The Pod keeps restarting.
Could Kubernetes be killing it because of a failed probe or memory issue?
I'll check the Pod status and events.Same commands.
Very different way of thinking.
What good debugging starts to look like
As you get better, your debugging process becomes less random.
Something breaks.
You ask:
What exactly failed?Then:
What changed?Then:
What's one possible reason?Then:
What can I check that would tell me whether that's true?You run the check.
Then you use the result to decide what comes next.
Over time, you'll naturally get faster because you've seen more failure patterns.
You'll recognise things like:
connection refused
DNS failures
expired certificates
missing environment variables
bad permissions
full disks
OOM kills
broken health checks
wrong routes
expired tokens
dependency timeouts
Experience helps.
But even experienced engineers don't know the answer immediately every time.
They're usually just better at narrowing the problem down.
The three questions I want you to remember
Good debugging isn't knowing 100 Linux commands.
It's being able to look at a broken system and keep making the problem smaller.
Whenever you're stuck, ask yourself:
1. What do I actually know?
2. What am I assuming?
3. What's the smallest thing I can check next?
Get good at those three questions and debugging, and I promise you’d climb up the career ladder much faster than others.
Join 1,000+ engineers becoming better DevOps & SRE professionals.
Every week, I share:
How I'd approach problems differently (real projects, real mistakes)
Career moves that actually work (not LinkedIn motivational posts)
Technical deep-dives that change how you think about infrastructure
No fluff. No roadmaps. Just what works when you're building real systems.

👋 Find me on Twitter | Linkedin | Connect 1:1
Thank you for supporting this newsletter.
Y’all are the best.
