← All posts

CPU throttling in Kubernetes: why your pod is slow even at 30% CPU

CPU throttling in Kubernetes is one of the most misunderstood performance problems in cloud-native infrastructure. Your pod can be throttled at 30% average CPU usage.

If you've ever seen an application respond slowly while the CPU dashboard shows modest utilization, CPU throttling in Kubernetes is likely the cause. It's one of the most counterintuitive problems in cloud-native infrastructure.

How CPU limits work in cgroups. When you set resources.limits.cpu: "500m" in a Kubernetes pod spec, the kernel's CFS (Completely Fair Scheduler) enforces this via cgroups. Every 100ms period, your container is allowed to use 50ms of CPU (500 millicores = 50% of one core). If your process uses its 50ms in the first 30ms of the period, it gets throttled for the remaining 70ms — regardless of whether the node has idle CPU available.

The throttle trap. A container that spikes briefly to handle a request can use its entire quota in a burst, then be throttled while other requests queue up. Average CPU usage looks fine at 30%, but p99 latency is terrible because every burst triggers a 70ms forced pause.

How to check if you're being throttled. Look at container_cpu_cfs_throttled_seconds_total and container_cpu_cfs_periods_total in Prometheus. The ratio of throttled periods to total periods is your throttle percentage. Anything above 5-10% is impacting latency.

Fixes.

*Raise the CPU limit.* The simplest fix. If the node has headroom, allowing the burst costs nothing at low average utilization.

*Remove the CPU limit entirely.* Controversial but effective. On a cluster with good bin-packing, removing limits lets pods use spare capacity during bursts. The risk is a noisy-neighbor pod starving others during peak load.

*Use a Burstable QoS class intentionally.* Set requests without limits. The pod gets guaranteed CPU up to the request amount and can burst beyond it when capacity is available.

*Tune the CFS period.* Lowering cpu.cfs_period_us from 100ms to 10ms gives finer-grained scheduling and smaller throttle windows. Available in newer kernels with cgroup v2.

The cpum.ai agent checks container_cpu_cfs_throttled_seconds_total when diagnosing slow Kubernetes workloads, and cross-references it with p99 latency from traces to confirm throttling is the latency cause rather than application logic.

See what's causing your CPU

cpum.ai turns CPU, process, disk, and memory signals into plain-English explanations with evidence.

Open cpum.ai
CPU throttling in Kubernetes: why your pod is slow even at 30% CPU — cpum.ai blog | cpum.ai