A flame graph collapses a CPU profile into a visual where the width of each frame is proportional to the time spent in that function. Reading one correctly takes about 30 seconds. Generating one on a live server takes about two minutes.
What you need. Linux perf (usually in linux-tools-$(uname -r) or the perf package), and Brendan Gregg's FlameGraph scripts from GitHub. For a non-root user: sudo perf, or adjust /proc/sys/kernel/perf_event_paranoid to 1.
Generating the profile. Record 30 seconds of CPU samples for a specific PID using: perf record -F 99 -p PID -g -- sleep 30. Then convert to a flamegraph with: perf script | stackcollapse-perf.pl | flamegraph.pl > flame.svg. Open the SVG in a browser — it's interactive.
How to read it. The x-axis is not time — it's alphabetically sorted stack traces. Width equals time spent. Look for wide frames near the top. If a frame spans 40% of the graph's width, 40% of CPU samples included that function in the call stack.
What to look for in production.
*Wide tower with a narrow top*: CPU is spent in one deep call path. The widest frame at the top is your hot function.
*Many equal-width columns*: work is well-distributed across many code paths. Harder to optimize without application-level redesign.
*[unknown] frames*: missing debug symbols. For interpreted languages (Python, Ruby, JVM), install language-specific perf extensions or use async-profiler (Java) or py-spy (Python) instead.
For Node.js. Run Node with --node-options='--perf-basic-prof' to generate readable JavaScript frames. Alternatively, clinic.js flame wraps this in a single command and generates an HTML report.
For Go. Use pprof: import the net/http/pprof package and hit /debug/pprof/profile?seconds=30. Visualize with: go tool pprof -http=:8080 cpu.prof.
The cpum.ai agent stores a 15-minute rolling window of per-PID CPU time. When you request a deep investigation, it identifies the top PIDs by CPU delta and recommends the appropriate profiling tool based on the process name and language runtime it detects.