If you've ever seen an application respond slowly while the CPU dashboard shows modest utilization, CPU throttling in Kubernetes is likely the cause. It's one of the most counterintuitive problems in cloud-native infrastructure.
How CPU limits work in cgroups. When you set resources.limits.cpu: "500m" in a Kubernetes pod spec, the kernel's CFS (Completely Fair Scheduler) enforces this via cgroups. Every 100ms period, your container is allowed to use 50ms of CPU (500 millicores = 50% of one core). If your process uses its 50ms in the first 30ms of the period, it gets throttled for the remaining 70ms — regardless of whether the node has idle CPU available.
The throttle trap. A container that spikes briefly to handle a request can use its entire quota in a burst, then be throttled while other requests queue up. Average CPU usage looks fine at 30%, but p99 latency is terrible because every burst triggers a 70ms forced pause.
How to check if you're being throttled. Look at container_cpu_cfs_throttled_seconds_total and container_cpu_cfs_periods_total in Prometheus. The ratio of throttled periods to total periods is your throttle percentage. Anything above 5-10% is impacting latency.
Fixes.
*Raise the CPU limit.* The simplest fix. If the node has headroom, allowing the burst costs nothing at low average utilization.
*Remove the CPU limit entirely.* Controversial but effective. On a cluster with good bin-packing, removing limits lets pods use spare capacity during bursts. The risk is a noisy-neighbor pod starving others during peak load.
*Use a Burstable QoS class intentionally.* Set requests without limits. The pod gets guaranteed CPU up to the request amount and can burst beyond it when capacity is available.
*Tune the CFS period.* Lowering cpu.cfs_period_us from 100ms to 10ms gives finer-grained scheduling and smaller throttle windows. Available in newer kernels with cgroup v2.
The cpum.ai agent checks container_cpu_cfs_throttled_seconds_total when diagnosing slow Kubernetes workloads, and cross-references it with p99 latency from traces to confirm throttling is the latency cause rather than application logic.