ποΈ 06082026 0000
Start here if a Grafana dashboard feels like a wall of unfamiliar numbers. Do not begin by memorizing metrics. Learn one path from the operating system to a panel, then reuse that method for CPU, storage, network, and applications.
The path is:
real system
β operating-system measurement
β exporter metric
β Prometheus samples
β PromQL calculation
β Grafana panel
β operational decision
First path: understand host memoryβ
Read these in order:
- how_linux_uses_memory β why βused memoryβ is not the same as βunavailable memoryβ
- linux_memory_pages β anonymous memory, file-backed memory, clean and dirty pages
- page_cache β why Linux spends spare RAM on faster file access
- proc_meminfo β where Linux reports
MemFree,Cached, slab, swap, andMemAvailable - interpreting_host_memory β how to recognize real host pressure
- host_memory_dashboard_walkthrough β how the Linux values become a Grafana panel and alert
- interpreting_container_memory β why a container can be under pressure even when the host is not, and vice versa
After this path, you should be able to explain:
- Why low
MemFreeis normal - Why
MemFree + Cachedis not a reliable available-memory formula - Why
MemAvailableis an estimate rather than a simple sum - Why historical swap usage is different from active swapping
- Why one memory panel cannot prove that an application leaks memory
- Why host and container memory alerts can disagree
Then learn how Prometheus represents the valueβ
Read:
- overview β the collection and query pipeline
- time_series_basics β samples, series, labels, and scrape intervals
- data_types β counters, gauges, histograms, and summaries
- range_function_calculations β rates and calculations over a time window
- node_exporter_host_metrics β lookup page for host metrics
Use these notes to explain a panel you already understand at the Linux level. Prometheus should answer βhow was the evidence collected and transformed?β, not replace the system model.
Read every dashboard panel with seven questionsβ
- Question β What operational question is the panel supposed to answer?
- Source β Which component measured the value?
- Metric β What does the raw metric mean, including its unit and type?
- Labels β What does each returned line represent?
- PromQL β What did the query aggregate, divide, or average?
- Display β How do the time range, unit, and threshold change the presentation?
- Conclusion β What can this panel prove, and what must be checked elsewhere?
If you cannot answer the first question, the panel may be decorative rather than diagnostic.
Continue by resourceβ
CPUβ
Read cpu_privilege_modes, interpreting_cpu_modes, cfs_bandwidth_control, and interpreting_cpu_usage. The important distinction is demand versus delivered CPU time versus time lost to throttling or another bottleneck.
Storageβ
Begin with inodes and the filesystem section of node_exporter_host_metrics. A future storage path should separate filesystem capacity from block-device performance.
Containersβ
Read linux_cgroups, cadvisor_container_metrics, interpreting_container_memory, and linux_oom_killer. A container is a cgroup accounting boundary, not a separate kernel.
Applicationsβ
Read prometheus_spring_boot_pipeline and visualization_queries after the foundation notes. Application panels are easier once counters, gauges, labels, and rates are familiar.
Use real dashboards as exercisesβ
For each section, choose one go2c panel and write down answers to the seven questions. Prefer panels tied to an alert or a real incident. Ignore unused exported metrics until a dashboard, alert, or investigation gives you a reason to learn them.