← All posts

Flame graphs and perf: a practical guide for production Linux servers

Flame graphs are the fastest way to understand where CPU time is actually going. Here's how to generate and read them on a live production server in under two minutes.

A flame graph collapses a CPU profile into a visual where the width of each frame is proportional to the time spent in that function. Reading one correctly takes about 30 seconds. Generating one on a live server takes about two minutes.

What you need. Linux perf (usually in linux-tools-$(uname -r) or the perf package), and Brendan Gregg's FlameGraph scripts from GitHub. For a non-root user: sudo perf, or adjust /proc/sys/kernel/perf_event_paranoid to 1.

Generating the profile. Record 30 seconds of CPU samples for a specific PID using: perf record -F 99 -p PID -g -- sleep 30. Then convert to a flamegraph with: perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg. Open the SVG in a browser — it's interactive.

How to read it. The x-axis is not time — it's alphabetically sorted stack traces. Width equals time spent. Look for wide frames near the top. If a frame spans 40% of the graph's width, 40% of CPU samples included that function in the call stack.

What to look for in production.

*Wide tower with a narrow top*: CPU is spent in one deep call path. The widest frame at the top is your hot function.

*Many equal-width columns*: work is well-distributed across many code paths. Harder to optimize without application-level redesign.

*[unknown] frames*: missing debug symbols. For interpreted languages (Python, Ruby, JVM), install language-specific perf extensions or use async-profiler (Java) or py-spy (Python) instead.

For Node.js. Run Node with --node-options='--perf-basic-prof' to generate readable JavaScript frames. Alternatively, clinic.js flame wraps this in a single command and generates an HTML report.

For Go. Use pprof: import the net/http/pprof package and hit /debug/pprof/profile?seconds=30. Visualize with: go tool pprof -http=:8080 cpu.prof.

The cpum.ai agent stores a 15-minute rolling window of per-PID CPU time. When you request a deep investigation, it identifies the top PIDs by CPU delta and recommends the appropriate profiling tool based on the process name and language runtime it detects.

See what's causing your CPU

cpum.ai turns CPU, process, disk, and memory signals into plain-English explanations with evidence.

Open cpum.ai
Flame graphs and perf: a practical guide for production Linux servers — cpum.ai blog | cpum.ai