Prometheus Alerting Rules That Do Not Wake You Up For Nothing

Prometheus Alerting Rules That Do Not Wake You Up For Nothing

Our monitoring guide covers collecting metrics. Collecting them is the easy part. Deciding what should wake someone at 3am is where monitoring projects usually go wrong, in both directions.

Alert on symptoms

The single most useful rule of thumb.

A cause: CPU is at 95 percent. Memory is nearly full. Disk I/O is saturated.

A symptom: requests are failing. Latency has tripled. The queue is not draining.

Alerting on causes produces noise, because a server at 95 percent CPU that serves every request in 20ms is working correctly. That is what you bought the CPU for.

Alerting on symptoms produces signal, because a request failing is a user having a bad time regardless of which resource is responsible.

Causes belong on dashboards, where they are how you diagnose an alert after it fires.

The for clause

groups:
  - name: availability
    interval: 30s
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
          /
          sum(rate(http_requests_total[5m])) by (service)
          > 0.05
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.service }} error rate is {{ $value | humanizePercentage }}"
          runbook: "https://wiki.example.com/runbooks/high-error-rate"

for: 10m is what makes this usable. The condition must hold for ten minutes before the alert fires. A brief spike during a deploy resolves itself and nobody is woken.

The cost is ten minutes of delay on a genuine incident. That is almost always the right trade, and the exception is something where ten minutes of downtime is unacceptable, in which case shorten it deliberately rather than removing it.

Always include a runbook link. The person receiving the alert at 3am may not be the person who wrote it. An alert that says what is wrong and not what to do about it is half an alert.

Rules worth having

      - alert: TargetDown
        expr: up == 0
        for: 5m
        labels: { severity: critical }
        annotations:
          summary: "{{ $labels.job }} on {{ $labels.instance }} is not responding"

      - alert: DiskWillFillIn4Hours
        expr: |
          predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4*3600) < 0
        for: 30m
        labels: { severity: warning }
        annotations:
          summary: "{{ $labels.instance }} {{ $labels.mountpoint }} fills in about 4 hours"

      - alert: HighLatency
        expr: |
          histogram_quantile(0.99,
            sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
          ) > 2
        for: 10m
        labels: { severity: warning }

      - alert: CertificateExpiringSoon
        expr: probe_ssl_earliest_cert_expiry - time() < 14 * 24 * 3600
        for: 1h
        labels: { severity: warning }

predict_linear is the one people do not know about and should. Rather than alerting when a disk is 90 percent full, which on a large disk may be fine for months, it alerts when the trend says it will fill within four hours. A disk at 60 percent filling rapidly is the urgent case; a disk at 92 percent that has been stable for a year is not.

Note fstype!~"tmpfs|overlay" to exclude container and tmpfs mounts, which otherwise generate constant noise.

Certificate expiry at fourteen days gives time to act. Our Let’s Encrypt guide covers automated renewal, and this alert is the safety net for when renewal silently stops working, which is the actual failure mode.

Recording rules

An expensive expression evaluated on every rule evaluation and in every dashboard panel is wasted work.

  - name: recording
    interval: 30s
    rules:
      - record: service:request_error_rate:ratio5m
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
          /
          sum(rate(http_requests_total[5m])) by (service)
      - alert: HighErrorRate
        expr: service:request_error_rate:ratio5m > 0.05
        for: 10m

The convention level:metric:operation in the record name is worth following, because it tells a reader what aggregation has already been applied.

Alertmanager

Rules decide when. Alertmanager decides who, how, and how often.

route:
  group_by: ['alertname', 'cluster', 'service']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  receiver: chat

  routes:
    - matchers: [ severity="critical" ]
      receiver: pager
      continue: false

    - matchers: [ severity="warning" ]
      receiver: chat

inhibit_rules:
  - source_matchers: [ alertname="TargetDown" ]
    target_matchers: [ severity="warning" ]
    equal: ['instance']

receivers:
  - name: pager
    webhook_configs:
      - url: 'https://events.example.com/prometheus'

  - name: chat
    webhook_configs:
      - url: 'https://chat.example.com/hooks/alerts'

Three mechanisms matter here.

Grouping. group_by collects related alerts into one notification. A rack losing power fires forty TargetDown alerts; grouping by cluster sends one message listing them.

group_wait holds the first notification briefly so related alerts arriving within thirty seconds are batched together.

Inhibition. The rule above suppresses warnings for an instance that is already known to be down. Without it, a dead host pages you once for being dead and then again for every service on it that stopped responding.

repeat_interval: 4h re-notifies about a still-firing alert. Too short and it becomes noise; too long and a forgotten incident stays forgotten.

Silences

amtool silence add alertname=HighLatency service=checkout \
  --duration=2h --comment="deploying, expected"
amtool silence query
amtool alert query

Silence before planned work. An alert that fires during a maintenance window you knew about is training people to ignore alerts.

Testing rules

promtool check rules alerts.yml
promtool test rules tests.yml
# tests.yml
rule_files:
  - alerts.yml
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      - series: 'up{job="api", instance="web01"}'
        values: '1+0x5 0+0x10'
    alert_rule_test:
      - eval_time: 12m
        alertname: TargetDown
        exp_alerts:
          - exp_labels:
              severity: critical
              job: api
              instance: web01

Unit tests for alert rules. This catches the classic mistake of an expression that never fires because of a label mismatch, which you would otherwise discover during an incident when it stayed silent.

The discipline

Few alerts. Every alert should mean someone does something. If the answer is “nothing, it clears up,” delete it.

Delete what nobody acts on. An alert that fires weekly and is routinely ignored is worse than no alert, because it trains people to ignore the channel where the real one arrives.

Review after incidents. Did an alert fire? Was it useful? Did something fire that was not useful? Adjust.

Separate urgent from informational. Paging is for things needing a human now. Everything else goes to a channel reviewed during working hours.

The failure mode to avoid is a monitoring system that produces so much noise people mute it, at which point you have the operational cost of monitoring and none of the benefit.

Frequently Asked Questions

What does the for clause do in a Prometheus alert rule?

It requires the condition to stay true for that duration before the alert fires. Without it, a single scrape where a value crossed a threshold sends a notification, so a brief spike wakes someone. A for clause of five or ten minutes removes most false alarms at the cost of delaying genuine ones by the same amount.

Should I alert on CPU usage?

Usually not. High CPU is a cause, not a symptom, and a server at 95 percent that serves every request quickly is working correctly. Alert on what users experience, such as error rates and latency, and use CPU as a dashboard metric for diagnosis after an alert fires.

What is the difference between grouping and inhibition in Alertmanager?

Grouping collects related alerts into a single notification so one incident produces one message. Inhibition suppresses alerts entirely when a more severe related alert is already firing, so a host being down does not also page you about every service on it.

How do I avoid alert fatigue?

Alert on few things, make every alert actionable, and delete alerts nobody acts on. An alert that fires regularly and is routinely ignored is worse than no alert, because it trains people to ignore the channel where real alerts arrive.

What is the difference between recording rules and alerting rules?

A recording rule precomputes an expression and stores the result as a new time series, which makes expensive queries fast. An alerting rule evaluates a condition and fires when true. Complex alerts frequently use a recording rule underneath so the evaluation stays cheap.

Should alerts go to email or chat?

Anything urgent needs a channel someone will notice within minutes, which usually means a paging service rather than email. Reserve paging for things requiring immediate human action, and send everything else to a chat channel or a ticket queue that gets reviewed rather than watched.