πŸ—“οΈ 06082026 0000

PROMETHEUS LEARNING PATH

Start here if a Grafana dashboard feels like a wall of unfamiliar numbers. Do not begin by memorizing metrics. Learn one path from the operating system to a panel, then reuse that method for CPU, storage, network, and applications.

The path is:

real system
β†’ operating-system measurement
β†’ exporter metric
β†’ Prometheus samples
β†’ PromQL calculation
β†’ Grafana panel
β†’ operational decision

First path: understand host memory​

Read these in order:

  1. how_linux_uses_memory β€” why β€œused memory” is not the same as β€œunavailable memory”
  2. linux_memory_pages β€” anonymous memory, file-backed memory, clean and dirty pages
  3. page_cache β€” why Linux spends spare RAM on faster file access
  4. proc_meminfo β€” where Linux reports MemFree, Cached, slab, swap, and MemAvailable
  5. interpreting_host_memory β€” how to recognize real host pressure
  6. host_memory_dashboard_walkthrough β€” how the Linux values become a Grafana panel and alert
  7. interpreting_container_memory β€” why a container can be under pressure even when the host is not, and vice versa

After this path, you should be able to explain:

  • Why low MemFree is normal
  • Why MemFree + Cached is not a reliable available-memory formula
  • Why MemAvailable is an estimate rather than a simple sum
  • Why historical swap usage is different from active swapping
  • Why one memory panel cannot prove that an application leaks memory
  • Why host and container memory alerts can disagree

Then learn how Prometheus represents the value​

Read:

  1. overview β€” the collection and query pipeline
  2. time_series_basics β€” samples, series, labels, and scrape intervals
  3. data_types β€” counters, gauges, histograms, and summaries
  4. range_function_calculations β€” rates and calculations over a time window
  5. node_exporter_host_metrics β€” lookup page for host metrics

Use these notes to explain a panel you already understand at the Linux level. Prometheus should answer β€œhow was the evidence collected and transformed?”, not replace the system model.

Read every dashboard panel with seven questions​

  1. Question β€” What operational question is the panel supposed to answer?
  2. Source β€” Which component measured the value?
  3. Metric β€” What does the raw metric mean, including its unit and type?
  4. Labels β€” What does each returned line represent?
  5. PromQL β€” What did the query aggregate, divide, or average?
  6. Display β€” How do the time range, unit, and threshold change the presentation?
  7. Conclusion β€” What can this panel prove, and what must be checked elsewhere?

If you cannot answer the first question, the panel may be decorative rather than diagnostic.

Continue by resource​

CPU​

Read cpu_privilege_modes, interpreting_cpu_modes, cfs_bandwidth_control, and interpreting_cpu_usage. The important distinction is demand versus delivered CPU time versus time lost to throttling or another bottleneck.

Storage​

Begin with inodes and the filesystem section of node_exporter_host_metrics. A future storage path should separate filesystem capacity from block-device performance.

Containers​

Read linux_cgroups, cadvisor_container_metrics, interpreting_container_memory, and linux_oom_killer. A container is a cgroup accounting boundary, not a separate kernel.

Applications​

Read prometheus_spring_boot_pipeline and visualization_queries after the foundation notes. Application panels are easier once counters, gauges, labels, and rates are familiar.

Use real dashboards as exercises​

For each section, choose one go2c panel and write down answers to the seven questions. Prefer panels tied to an alert or a real incident. Ignore unused exported metrics until a dashboard, alert, or investigation gives you a reason to learn them.

References​