Every week the cpum.ai investigation agent runs on hundreds of Linux servers. Most spikes are unsurprising — a deploy, a noisy neighbour, a runaway query. Some are not. Here are five that stood out.
1. The kworker that wasn't idle. A user reported 60% CPU on a host with no foreground load. Top showed only kernel workers. Cause: a stuck NFS mount triggering retransmit storms in soft-IRQ context. Fix: unmount + remount with a sane timeout.
2. Python script using 100% of a single core for 9 hours straight. Looked like an infinite loop. Was actually JSON.parse on a 4 GB log file with a regex that backtracked exponentially.
3. Postgres autovacuum at 04:00 UTC. Hourly autovacuum on a hot table generated read-IO bursts that pushed iowait above 40% — small enough to miss, big enough to push p99 latency past SLO.
4. Container CPU throttling that wasn't. A pod's CPU graph showed 100% sustained but throttling counters were zero. Cause: cgroup v1 vs v2 interpretation mismatch in the metrics collector.
5. CPU steal time spike on a "dedicated" instance. Customer was on AWS m5.large but saw 12% steal. Hypervisor scheduling on the underlying host had drifted. Migrated to a c6i.large and steal returned to <0.5%.
Want your weirdest spike featured next week? Publish your diagnosis from the dashboard and tag it with "weekly".