Blog

Notes from the diagnostic floor

Weekly incident roundups, observability essays, and plain-English performance writing.

cpulinuxperformancehardware

CPU frequency scaling and performance: why your server isn't running at full speed

Modern CPUs don't run at their rated speed all the time. Understanding frequency scaling, turbo boost, and power governors can unlock hidden performance on Linux servers.

Read post →
ebpfprofilinglinuxobservability

eBPF for CPU profiling: what it is, when to use it, and how to start

eBPF lets you profile any process, including the kernel itself, with near-zero overhead and no restarts. Here's a practical introduction for production diagnostics.

Read post →
perfprofilinglinuxflame-graphs

Flame graphs and perf: a practical guide for production Linux servers

Flame graphs are the fastest way to understand where CPU time is actually going. Here's how to generate and read them on a live production server in under two minutes.

Read post →
kubernetescgroupscpucloud

CPU throttling in Kubernetes: why your pod is slow even at 30% CPU

CPU throttling in Kubernetes is one of the most misunderstood performance problems in cloud-native infrastructure. Your pod can be throttled at 30% average CPU usage.

Read post →
postgresdatabasecputriage

Diagnosing Postgres CPU spikes: the five-minute triage

PostgreSQL CPU spikes are almost always caused by one of five things. Here's how to identify which one in under five minutes using queries you already have access to.

Read post →
linuxkernelcpudiagnostics

Kernel time vs user time: what the split tells you about your system

The us/sy split in top is one of the most underused diagnostics available. A high system time percentage almost always points to a specific class of problem.

Read post →
nodejscpuprofiling

Why your Node.js server spikes to 100% on a single core

Node.js is single-threaded by design. When CPU pegs one core, it's almost always the event loop. Here's how to find the exact function causing it.

Read post →
linuxinternalscpu

Reading /proc/stat correctly: the source of truth for CPU metrics

Every CPU percentage you've ever seen ultimately comes from /proc/stat. Understanding what it actually measures — and what it doesn't — makes you a better diagnostician.

Read post →
incident-roundupweeklydiagnostics

Five weirdest CPU spikes we saw this week

From a kworker eating 60% CPU on a freshly-rebooted host to a Python script exhausting file descriptors at 02:14 UTC — five real diagnostics from the past seven days.

Read post →
philosophyobservability

Why "CPU is at 87%" is the wrong question

Most monitoring tools show you what your CPU number is. cpum.ai shows you why. Here's the difference, and why it matters at 02:14 in the morning.

Read post →
Blog — cpum.ai | cpum.ai