Blog
Notes from the diagnostic floor
Weekly incident roundups, observability essays, and plain-English performance writing.
CPU frequency scaling and performance: why your server isn't running at full speed
Modern CPUs don't run at their rated speed all the time. Understanding frequency scaling, turbo boost, and power governors can unlock hidden performance on Linux servers.
Read post →eBPF for CPU profiling: what it is, when to use it, and how to start
eBPF lets you profile any process, including the kernel itself, with near-zero overhead and no restarts. Here's a practical introduction for production diagnostics.
Read post →Flame graphs and perf: a practical guide for production Linux servers
Flame graphs are the fastest way to understand where CPU time is actually going. Here's how to generate and read them on a live production server in under two minutes.
Read post →CPU throttling in Kubernetes: why your pod is slow even at 30% CPU
CPU throttling in Kubernetes is one of the most misunderstood performance problems in cloud-native infrastructure. Your pod can be throttled at 30% average CPU usage.
Read post →Diagnosing Postgres CPU spikes: the five-minute triage
PostgreSQL CPU spikes are almost always caused by one of five things. Here's how to identify which one in under five minutes using queries you already have access to.
Read post →Kernel time vs user time: what the split tells you about your system
The us/sy split in top is one of the most underused diagnostics available. A high system time percentage almost always points to a specific class of problem.
Read post →Why your Node.js server spikes to 100% on a single core
Node.js is single-threaded by design. When CPU pegs one core, it's almost always the event loop. Here's how to find the exact function causing it.
Read post →Reading /proc/stat correctly: the source of truth for CPU metrics
Every CPU percentage you've ever seen ultimately comes from /proc/stat. Understanding what it actually measures — and what it doesn't — makes you a better diagnostician.
Read post →Five weirdest CPU spikes we saw this week
From a kworker eating 60% CPU on a freshly-rebooted host to a Python script exhausting file descriptors at 02:14 UTC — five real diagnostics from the past seven days.
Read post →Why "CPU is at 87%" is the wrong question
Most monitoring tools show you what your CPU number is. cpum.ai shows you why. Here's the difference, and why it matters at 02:14 in the morning.
Read post →