🗓️ 29052026 1500

INTERPRETING CPU MODES

node_cpu_seconds_total has a mode label that breaks CPU time into categories: user, system, iowait, steal, irq, softirq, nice, idle. The elevated mode tells you why CPU is busy — and whether adding cores, fixing code, or changing infrastructure is the right response.

Container-level CPU (see cadvisor_container_metrics) does not expose mode breakdown. You need node exporter for this. For the raw metric and queries, see node_exporter_host_metrics.

What Each Mode Means​

ModeCPU is doingElevated means
userRunning application codeApp is compute-bound
systemRunning kernel code (syscalls)Heavy I/O ops, context switching, network stack work
iowaitIdle, waiting for disk I/ODisk is the bottleneck, not CPU
stealHypervisor gave this slice to another VMVM is preempted by neighbors
irqHandling hardware interruptsRare; hardware issue or misconfigured NIC
softirqHandling software interrupts (packets, timers)High network packet rate
niceRunning low-priority user processesBackground jobs consuming CPU
idleDoing nothingCPU has capacity

Diagnosing by Mode Pattern​

High User, Low System​

  • Normal application workload — CPU is doing useful compute
  • Scale horizontally, optimize hot code paths, or accept the load

High System​

  • Kernel is working hard on behalf of applications
  • Common causes: excessive small I/O operations, heavy logging to disk, many short-lived connections, frequent context switches
  • Look at syscall patterns, not just application code

High Iowait​

  • CPU is idle but blocked waiting for disk. Adding more CPU cores will not help.
  • Common causes: slow disk (HDD vs SSD), sequential I/O on a busy volume, database queries scanning large tables
  • Check disk throughput and IOPS via node_disk_* metrics
  • Caveat: iowait only counts CPUs that are idle and waiting. If other work fills those CPUs, iowait drops even though disk is still slow

High Steal​

  • The hypervisor is taking CPU time for other tenants. No application-level fix.
  • Resize the instance, move to dedicated hosts, or talk to the cloud provider
  • > 5% sustained = worth investigating. > 10% = performance-impacting

High Softirq​

  • Software interrupt processing, usually network packet handling
  • Common on load balancers, reverse proxies, services handling thousands of small requests/sec
  • Consider kernel tuning (RPS/RFS) or scaling out

Typical Baselines by Workload​

Workloadusersystemiowaitsteal
Web server40–60%5–15%< 5%0%
Database20–40%10–20%10–30% (SSDs)0%
Batch / ML training70–95%2–5%< 2%0%
Load balancer10–30%10–20%< 1%0%
WARNING

High iowait with low overall CPU utilization is deceptive. The node looks "idle" but applications are stalled on disk. Check disk metrics before concluding the node has spare capacity.

Why This Matters for Containers​

  • container_cpu_usage_seconds_total shows total CPU consumed but not what kind of work
  • A container using 2 cores could be compute-bound (user), syscall-heavy (system), or stuck on I/O (iowait at the node level)
  • When container CPU looks high or latency spikes, check node CPU modes for the cause
  • If the node shows high steal, all containers on that node are affected regardless of their individual metrics

References​