Every Tuesday - Deep dives, architecture lessons, and real engineering stories.

Every Saturday - The best DevOps, SRE, Cloud, AI, tools, tutorials, and projects from the week.

📬 Missed This Week’s Uptime Sync?

This week’s best DevOps, SRE, Cloud, Kubernetes, database, AI infrastructure, and production engineering reads:

  • How Cloudflare saved another 100 TB of RAM with math and Rust

  • How Linux readahead works—and when to change it

  • Why a Postgres migration passed every test but failed with real data

  • Why every AI agent may need its own computer

  • How n8n was deployed, maintained, and migrated on EKS

  • PostgreSQL row-level security performance in practice

  • How an OpenAI DNS jailbreak actually worked

Plus: debugging stuck Kubernetes rollouts, understanding environment variables, using Claude Code from anywhere, scaling Docker on Railway, managing node drains with Pod Disruption Budgets, and tools for SQLite, AI security, OpenTofu, dev containers, and QUIC.

I’m thinking of creating practical DevOps/SRE interview guides for Kubernetes, Cloud, Linux, SRE, DevOps, and more.

Would you prefer buying:

Login or Subscribe to participate

I know, I know.

Debugging is a boring topic to read.

I could have written about whatever is trending this week - JEV, Gemini 4, some new AI model, agent framework, or another shiny tool everyone is talking about.

That probably would've gotten more clicks.

But if you're trying to become a genuinely good engineer, I'd rather you understand this.

Because once you get good at debugging, the real world becomes much easier.

Tools will change. Frameworks will change. The stack you use today might be completely different three years from now.

But production will still break.

An API will suddenly start returning 401.
A deployment that worked yesterday will start throwing 500.
A container will be running perfectly fine but somehow have no logs.
A script will fail with an error that makes absolutely no sense at first.

And in those moments, knowing another Kubernetes command or memorising another system design diagram won't save you.

How you think will.

Start with what you actually know

Imagine someone tells you:

❝

The app is broken.

That doesn't give you much to work with.

Instead, write down what is actually happening.

Request: GET /profile

Expected: 200 OK with profile JSON

Actual: 401 Unauthorized

Scope: happens in my local test client only

Started: after I changed the client environment variables

Recent change: renamed API_TOKEN to AUTH_TOKEN

Now you have something you can investigate.

This distinction matters.

A 401 Unauthorized is something you know happened.

❝

The authentication config is broken.

That's only a guess.

Other possible reasons could be:

  • the Authorization header was never sent

  • the token expired

  • you're calling the wrong endpoint

  • the server is reading a different environment variable

  • a proxy changed the request

Debugging gets much easier when you stop treating your first guess like the answer.

For a small script, the same idea applies.

Instead of:

❝

My Python script doesn't work.

Write:

Command: python totals.py

Expected: total is 15

Actual: TypeError: unsupported operand type(s)

Input: quantity came from a text file

You now have a specific failure that you can reproduce.

That's already a much better starting point.

When something breaks, do this

You don't need a complicated debugging framework.

Use these five steps.

1. What exactly is broken?

Be specific.

Bad:

The API is broken.

Better:

GET /profile should return 200 but returns 401.

Bad:

Docker logs aren't working.

Better:

The container is running, the request fails, but docker logs returns nothing.

The more clearly you describe the problem, the easier it becomes to decide what to check next.

2. What changed?

A lot of bugs appear shortly after something changed.

Check things like:

  • a new deployment

  • configuration changes

  • environment variables

  • dependency upgrades

  • different input

  • DNS changes

  • certificates

  • permissions

  • network rules

  • infrastructure changes

This doesn't mean the latest change is automatically the cause.

It just gives you somewhere useful to start looking.

3. Pick one possible reason

Don't try to explain the entire problem at once.

Pick one thing that could explain what you're seeing.

For example:

Maybe the Authorization header isn't being sent.

That's enough.

4. Find one way to check it

Now ask:

❝

If this is actually the problem, what should I see?

If the client isn't sending an Authorization header, inspecting the request should show that the header is missing.

So run a check that answers exactly that question.

5. Let the result choose your next step

This is where good debugging starts to feel very different from guessing.

If the header is missing:
> investigate the client.

If the header exists:
> stop touching the client and move to the server.

Each result should make your search area smaller. That's the goal.

Example 1: API returns 401 Unauthorized

Suppose your local API should return:

200 OK

but you're getting:

401 Unauthorized

A random approach might look like this:

  • edit authentication code

  • restart the server

  • change an environment variable

  • disable an auth check

  • restart everything

  • try again

Maybe it starts working. But which change fixed it? You don't know.

And worse, maybe disabling the auth check was the thing that made it work.

That's not really a fix.

Instead, start with one possible reason:

Maybe the client isn't sending the Authorization header.

Now ask:

If that's true, what should I see?

You should see that the outgoing request doesn't contain an Authorization header.

So inspect the request:

curl -v http://localhost:8080/profile

With curl -v, outgoing request lines start with > and response lines start with <.

Look for something like:

> GET /profile HTTP/1.1
> Host: localhost:8080
> Authorization: Bearer ...

If the header isn't there, you just narrowed the problem down.

Now check the client.

Maybe your code expects:

API_TOKEN

but your environment only contains:

AUTH_TOKEN

That's a useful finding.

You went from:

Authentication is broken.

to:

The client isn't sending the token because it's reading the wrong environment variable.

Much better.

But what if the header is present? Good.

Your original guess was wrong. Stop changing client code. Move to the server.

Check:

  • server authentication logs

  • token validation

  • token expiry

  • environment variables

  • middleware

  • proxy configuration

A wrong guess is still useful when it tells you where not to look.

One thing to remember: curl -v shows what the client tried to send. If there is a proxy between the client and server, the request could still be changed later.

Also be careful before sharing verbose output. It can contain tokens, cookies, and other secrets.

Example 2: docker logs is empty

Here's another common one.

Your container is running.

Requests to the application are failing.

So you run:

docker logs my-app

And get nothing useful back.

It's easy to conclude:

❝

The application didn't log any errors.

But that's not necessarily true.

A better question is:

❝

Where is this application actually writing its logs?

docker logs normally shows output Docker captured from the container process, usually stdout and stderr.

But an application might instead write logs to:

/var/log/app/error.log

or send them somewhere else entirely.

So start with one possible reason:

Maybe the application writes errors to a file instead of stdout or stderr.

Now inspect what Docker actually started:

docker inspect my-app --format '{{.Path}} {{json .Args}}'

If the image has a shell, you can inspect the container:

docker exec -it my-app sh

Then check the application's configuration.

Look for:

  • log file paths

  • logging configuration

  • process arguments

  • environment variables

  • external logging agents

Maybe you discover:

log_file=/var/log/app/error.log

Now you know why docker logs looked empty.

The application was logging. Docker just wasn't showing those logs.

The exact commands will differ between applications and container images.

Some containers don't even include a shell. That's fine.

The useful part is the order:

> docker logs is empty
> Where does the app write logs?
> What process is actually running?
> What logging configuration is it using?
> Where should I look next?

That's debugging.

Don't confuse fixing the symptom with finding the cause

This becomes especially important in production.

Imagine you deploy a new release.

A few minutes later:

POST /checkout

starts returning:

500 Internal Server Error

Customers are affected. You roll the release back. The error rate drops. Users can check out again.

But what did you actually learn?

You know:

Rollback changed the system and the errors stopped.

You do not automatically know:

The new application code caused the problem.

The rollback may also have changed:

  • traffic patterns

  • container instances

  • cache state

  • connections

  • configuration

  • dependency behaviour

  • timing

So keep two separate questions in your head.

Question 1: How do we stop users from being affected?

That might mean:

  • rollback

  • disable a feature

  • route traffic elsewhere

  • scale something up

  • fail over

  • restart unhealthy instances

Question 2: Why did this happen?

That investigation can continue after the system is stable.

During an incident, write down what you actually know:

Impact:
Checkout requests return 500 for some users

Start time:
14:07 UTC

Scope:
One region

Mitigation:
Rolled back release 2025.03.18.2

Observed result:
Error rate returned to normal

Root cause:
Unknown

Notice the last line.

Root cause: Unknown

There's nothing wrong with that.

It's much better than pretending your first theory is already confirmed. This is why incident response and root-cause analysis are different jobs.

During an outage, your first priority may be getting the service healthy again.

Understanding exactly why it happened can come next.

A command should answer a question

One habit I see beginners develop is collecting debugging commands.

top
free -m
df -h
netstat
docker logs
kubectl describe
journalctl

These are useful tools.

But knowing commands isn't the same thing as knowing how to debug.

Before running one, ask:

❝

What am I trying to learn from this?

Instead of:

Let me run top.

Think:

The service became slow.

Could CPU saturation explain it?

I'll check CPU usage.

Instead of:

Let me run df -h.

Think:

The application can't write new files.

Could the disk be full?

I'll check filesystem usage.

Instead of:

Let me look at Kubernetes events.

Think:

The Pod keeps restarting.

Could Kubernetes be killing it because of a failed probe or memory issue?

I'll check the Pod status and events.

Same commands.

Very different way of thinking.

What good debugging starts to look like

As you get better, your debugging process becomes less random.

Something breaks.

You ask:

What exactly failed?

Then:

What changed?

Then:

What's one possible reason?

Then:

What can I check that would tell me whether that's true?

You run the check.

Then you use the result to decide what comes next.

Over time, you'll naturally get faster because you've seen more failure patterns.

You'll recognise things like:

  • connection refused

  • DNS failures

  • expired certificates

  • missing environment variables

  • bad permissions

  • full disks

  • OOM kills

  • broken health checks

  • wrong routes

  • expired tokens

  • dependency timeouts

Experience helps.

But even experienced engineers don't know the answer immediately every time.

They're usually just better at narrowing the problem down.

The three questions I want you to remember

Good debugging isn't knowing 100 Linux commands.

It's being able to look at a broken system and keep making the problem smaller.

Whenever you're stuck, ask yourself:

1. What do I actually know?
2. What am I assuming?
3. What's the smallest thing I can check next?

Get good at those three questions and debugging, and I promise you’d climb up the career ladder much faster than others.

Join 1,000+ engineers becoming better DevOps & SRE professionals.

Every week, I share:

  • How I'd approach problems differently (real projects, real mistakes)

  • Career moves that actually work (not LinkedIn motivational posts)

  • Technical deep-dives that change how you think about infrastructure

No fluff. No roadmaps. Just what works when you're building real systems.

👋 Find me on Twitter | Linkedin | Connect 1:1

Thank you for supporting this newsletter.
Y’all are the best.

Reply

Avatar

or to participate