🗓️ 18052026 2200

NODE EXPORTER HOST METRICS

The Prometheus node exporter exposes hardware and OS-level metrics per host — CPU, memory, disk, network, filesystem — all prefixed with node_. It answers "what does the machine have and how much is used?"

cadvisor_container_metrics tells you what containers consume; node exporter tells you what the underlying host has available. Together: "is the container starved, or is the host overloaded?"

How Node Exporter Works​

  • Runs as a DaemonSet in K8s (one per node) or standalone binary
  • Endpoint: :9100/metrics
  • Key label: instance (identifies the node)
  • Reads from /proc and /sys — zero app instrumentation needed

CPU Metrics​

Single metric, multiple modes via the mode label:

MetricTypeUnit
node_cpu_seconds_totalCounterseconds

Modes: user, system, idle, iowait, irq, softirq, steal, nice, guest

PromQL​

Overall CPU utilization %:

100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)

CPU breakdown by mode:

avg by(mode) (rate(node_cpu_seconds_total[5m])) * 100

iowait (disk bottleneck signal):

avg by(instance) (rate(node_cpu_seconds_total{mode="iowait"}[5m])) * 100

Number of CPU cores:

count by(instance) (count by(instance, cpu) (node_cpu_seconds_total))

Steal time (noisy neighbor in shared infra):

avg by(instance) (rate(node_cpu_seconds_total{mode="steal"}[5m])) * 100

Memory Metrics​

MetricTypeUnitMeasures
node_memory_MemTotal_bytesGaugebytesTotal physical RAM
node_memory_MemAvailable_bytesGaugebytesAvailable memory (kernel estimate, includes reclaimable)
node_memory_MemFree_bytesGaugebytesCompletely free memory (misleadingly low)
node_memory_Buffers_bytesGaugebytesBuffer cache
node_memory_Cached_bytesGaugebytesPage cache
node_memory_SwapTotal_bytesGaugebytesTotal swap
node_memory_SwapFree_bytesGaugebytesFree swap
WARNING

Use MemAvailable, not MemFree. Linux counts cache and buffers as "used", but they are reclaimable under pressure. MemAvailable accounts for this — MemFree makes nodes look worse than they are.

PromQL​

Memory utilization %:

(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100

Swap usage %:

(1 - node_memory_SwapFree_bytes / node_memory_SwapTotal_bytes) * 100

Available memory in GB:

node_memory_MemAvailable_bytes / 1024 / 1024 / 1024

Disk / Filesystem Metrics​

Filesystem (capacity)​

MetricTypeUnit
node_filesystem_size_bytesGaugebytes
node_filesystem_avail_bytesGaugebytes
node_filesystem_free_bytesGaugebytes

avail_bytes = space available to non-root users. free_bytes = total free including reserved blocks.

Disk I/O (throughput)​

MetricTypeUnit
node_disk_read_bytes_totalCounterbytes
node_disk_written_bytes_totalCounterbytes
node_disk_reads_completed_totalCounterops
node_disk_writes_completed_totalCounterops
node_disk_io_time_seconds_totalCounterseconds
node_disk_read_time_seconds_totalCounterseconds
node_disk_write_time_seconds_totalCounterseconds

PromQL​

Filesystem usage % (exclude tmpfs/overlay):

(1 - node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay"}) * 100

Disk I/O throughput (read + write bytes/sec):

rate(node_disk_read_bytes_total[5m])
+ rate(node_disk_written_bytes_total[5m])

Disk I/O utilization % (how busy the disk is):

rate(node_disk_io_time_seconds_total[5m]) * 100

IOPS:

rate(node_disk_reads_completed_total[5m])
+ rate(node_disk_writes_completed_total[5m])

Average I/O latency per write:

rate(node_disk_write_time_seconds_total[5m])
/ rate(node_disk_writes_completed_total[5m])

Predict disk full (capacity planning) — see range_function_calculations for how predict_linear works:

predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 24*3600) < 0

Network Metrics​

MetricTypeUnit
node_network_receive_bytes_totalCounterbytes
node_network_transmit_bytes_totalCounterbytes
node_network_receive_errs_totalCountererrors
node_network_transmit_errs_totalCountererrors
node_network_receive_drop_totalCounterpackets
node_network_transmit_drop_totalCounterpackets

PromQL​

Network bandwidth by interface (exclude loopback):

rate(node_network_receive_bytes_total{device!="lo"}[5m])

Total network errors:

rate(node_network_receive_errs_total[5m])
+ rate(node_network_transmit_errs_total[5m])

Dropped packets:

rate(node_network_receive_drop_total[5m])
+ rate(node_network_transmit_drop_total[5m])

System Metrics​

MetricMeasures
node_load1 / node_load5 / node_load151/5/15-minute load averages
node_boot_time_secondsUnix timestamp of last boot
node_time_secondsCurrent system clock
node_uname_infoKernel/OS info (labels only, value always 1)

PromQL​

Uptime:

time() - node_boot_time_seconds

Load per CPU core (>1.0 = overloaded):

node_load1
/ count by(instance) (count by(instance, cpu) (node_cpu_seconds_total))

Quick Reference​

┌──────────────────────────┬──────────────────────────────────────────────────────────────────────┐
│ I want to know… │ PromQL │
├──────────────────────────┼──────────────────────────────────────────────────────────────────────┤
│ CPU utilization % │ 100 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100 │
│ Memory utilization % │ (1 - MemAvailable / MemTotal) * 100 │
│ Disk space % │ (1 - avail_bytes / size_bytes) * 100 │
│ Disk I/O throughput │ rate(node_disk_read_bytes_total[5m]) + rate(..written..[5m]) │
│ Disk I/O utilization % │ rate(node_disk_io_time_seconds_total[5m]) * 100 │
│ Network bandwidth │ rate(node_network_receive_bytes_total{device!="lo"}[5m]) │
│ Uptime │ time() - node_boot_time_seconds │
│ Load per CPU │ node_load1 / count(cpus) │
│ Disk full prediction │ predict_linear(avail_bytes[6h], 86400) < 0 │
└──────────────────────────┴──────────────────────────────────────────────────────────────────────┘

References​