{
 "id": "monitoring-stack",
 "kind": "skill",
 "name": "Monitoring and logging",
 "description": "Monitoring and logging: Prometheus, node_exporter, Grafana, Icinga 2 and the Elastic stack. Use for alert rules, PromQL, dashboards, check configs, log pipelines.",
 "version": "1.0.0",
 "author": "Hexa Hub",
 "files": {
  "SKILL.md": "---\nname: monitoring-stack\ndescription: Monitoring and logging: Prometheus, node_exporter, Grafana, Icinga 2 and the Elastic stack. Use for alert rules, PromQL, dashboards, check configs, log pipelines.\ntitle: Monitoring and logging\nicon: tabler:chart-line\ncategory: Monitoring\n---\n\n# Monitoring stack\n\nVersions differ a lot (Icinga 2 vs Director, Grafana 9 vs 11, Elastic 7 vs 8). Ask the version if it matters, and check\nthe official docs (`devtools__web_search` topic `prometheus`, `grafana`, `icinga`, `elastic`).\n\n## Prometheus and node_exporter\n- Metric types: counter (only goes up; always wrap in `rate()`/`increase()`), gauge, histogram (use\n  `histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m])))`).\n- Validate: `promtool check config prometheus.yml`, `promtool check rules rules.yml`, `promtool test rules`.\n- Useful node_exporter metrics: `node_cpu_seconds_total` (`mode=\"idle\"`), `node_memory_MemAvailable_bytes`,\n  `node_filesystem_avail_bytes` (exclude `tmpfs`/`overlay`), `node_load1`, `node_network_receive_bytes_total`, `up`.\n- CPU used %: `100 - avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100`.\n- Alerts: use `for:` to avoid flapping, meaningful labels (`severity`), and annotations with the value and runbook link.\n  Watch label cardinality (never put user IDs/URLs in labels).\n\n## Grafana\n- Dashboards as JSON/provisioning in version control; data sources provisioned in YAML. Use variables for instance/job.\n- In queries use `$__rate_interval` for `rate()`. Set units and thresholds; keep panels few and readable.\n\n## Icinga 2\n- Config objects: Host, Service, CheckCommand, Notification, User, TimePeriod; use `apply Service ... for` rules and\n  groups instead of repeating objects. Validate with `icinga2 daemon -C` before `systemctl reload icinga2`.\n- Distributed setups use zones/endpoints and `icinga2 node wizard`; mind the certificates. Plugins live in the\n  monitoring-plugins path; test a check by running the plugin by hand first, and check its exit code\n  (0 OK, 1 WARNING, 2 CRITICAL, 3 UNKNOWN).\n\n## Elastic stack\n- Elasticsearch: define index templates/mappings (keyword vs text), use ILM for retention, avoid wildcard-heavy queries,\n  keep shard counts sane. Kibana: KQL for quick filters, Lucene/ES|QL for more. Shippers: Elastic Agent or Beats;\n  parse logs with ingest pipelines (grok/dissect). Secure with TLS and API keys; don't expose port 9200.\n"
 }
}
