Centralized Logging with Loki
Our lnav guide covers reading logs on one machine. This covers the point where you have twenty machines and correlating an incident means SSHing into each of them.
Loki is Grafana’s log aggregation system, and its defining choice is what it does not index.
The design decision
Elasticsearch builds a full-text inverted index over log content. Any search is fast, and the index is frequently larger than the logs themselves. That is why an ELK stack is expensive to run.
Loki indexes only labels, stores the log lines as compressed chunks, and searches content by scanning the chunks that match your label selector.
The consequences follow directly:
Storage is dramatically cheaper, frequently an order of magnitude against Elasticsearch for the same logs.
Ingestion is cheap, because there is almost nothing to index.
You must narrow by label first. A query with no label selector is asking Loki to scan everything, which is slow and which the system will refuse past a limit.
That last point is the whole adjustment. Elasticsearch lets you search for a string across everything; Loki wants you to say which streams first.
Labels and cardinality
This is where people break Loki, and it is worth getting right before deploying anything.
Each unique combination of label values creates a stream. Good labels are bounded:
{job="nginx", host="web01", env="prod"}
A handful of jobs, a few dozen hosts, three environments. A few hundred streams.
Bad labels are unbounded:
{job="nginx", request_id="8f3a...", user_id="41288"}
Every request creates a new stream. Millions of streams, an index that no longer fits in memory, and a system that falls over.
Put high-cardinality values in the log line, not in a label. You can still search for them; Loki scans for them rather than indexing them, which is exactly the trade it is built around.
This is the same cardinality discipline Prometheus requires, and for the same reason.
Deploying
services:
loki:
image: grafana/loki:3.3.0
command: -config.file=/etc/loki/config.yml
volumes:
- ./loki-config.yml:/etc/loki/config.yml:ro
- loki-data:/loki
networks: [internal]
restart: unless-stopped
alloy:
image: grafana/alloy:latest
command: run --server.http.listen-addr=0.0.0.0:12345 /etc/alloy/config.alloy
volumes:
- ./alloy-config.alloy:/etc/alloy/config.alloy:ro
- /var/log:/var/log:ro
- /var/lib/docker/containers:/var/lib/docker/containers:ro
- /run/systemd/journal:/run/systemd/journal:ro
networks: [internal]
restart: unless-stopped
grafana:
image: grafana/grafana:latest
ports:
- "127.0.0.1:3000:3000"
networks: [internal]
restart: unless-stopped
volumes:
loki-data:
networks:
internal:
internal: true
Grafana Alloy supersedes Promtail as the agent. One binary handling logs, metrics, and traces, rather than a separate collector for each.
Note the read-only mounts and the internal network, per our container security guide.
Reading the journal directly
Worth calling out because most guides show tailing files, and on a systemd host that is the wrong input.
loki.source.journal "system" {
forward_to = [loki.write.default.receiver]
labels = { job = "systemd-journal", host = "web01" }
max_age = "12h"
}
loki.write "default" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
}
}
Reading the journal natively gets you structured fields that file tailing throws away: the unit name, PID, UID, and priority are already there rather than needing regex extraction.
loki.relabel "journal" {
forward_to = [loki.write.default.receiver]
rule {
source_labels = ["__journal__systemd_unit"]
target_label = "unit"
}
rule {
source_labels = ["__journal_priority_keyword"]
target_label = "level"
}
}
Now {unit="nginx.service", level="err"} is a label selector rather than a text search.
LogQL
{job="nginx"}
{job="nginx"} |= "500"
{job="nginx"} |= "500" != "/healthz"
{job="nginx"} | json | status >= 500
{unit="sshd.service"} |~ "Failed password for .* from ([0-9.]+)"
|= contains, != does not contain, |~ regex match. Chain them, and put the cheapest filter first because each stage reduces what the next has to scan.
Metrics from logs
The feature that makes Loki worth running next to Prometheus:
rate({job="nginx"} |= "500" [5m])
sum by (host) (rate({job="nginx"} | json | status >= 500 [5m]))
quantile_over_time(0.99,
{job="myapp"} | json | unwrap duration_ms [5m]
) by (endpoint)
Those are Prometheus-shaped queries over log data, which means they graph in Grafana and alert through Grafana alerting exactly like a metric. A 99th-percentile latency from application logs, without the application exporting a metric.
Our Prometheus alerting guide covers the alerting side, and our monitoring overview covers the wider stack.
Retention
limits_config:
retention_period: 720h
ingestion_rate_mb: 16
max_streams_per_user: 10000
compactor:
retention_enabled: true
delete_request_store: filesystem
storage_config:
aws:
s3: s3://minio:9000/loki
s3forcepathstyle: true
max_streams_per_user is a guard rail against the cardinality mistake. Hitting it is a signal that a label is unbounded, and it is far better than discovering the problem when the system falls over.
Object storage, whether S3 or self-hosted MinIO, makes retention a lifecycle policy rather than a disk sizing exercise.
What it does not replace
Local logs remain authoritative. When the network is down, when Alloy is broken, when Loki itself is the thing failing, journalctl on the box is what you have. Shipping logs is for correlation across machines, not for deleting what is on disk.
It is not an audit log. Logs shipped over the network by an agent that can be stopped are not tamper-evident. Our auditd guide covers the tool for that requirement.
Single-machine investigation is still better locally. For one server and one incident, lnav on the box beats a web interface.
Loki earns its place at the point where “which of these twenty machines was involved” is a question you ask regularly. Below that, it is infrastructure you are maintaining for no gain.
Frequently Asked Questions
How is Loki different from Elasticsearch for logs?
Elasticsearch builds a full-text index of log content, which makes arbitrary searches fast and the index large and expensive. Loki indexes only labels and stores the log lines compressed, so storage costs a fraction and searching content means scanning the selected streams rather than consulting an index.
What is a label and why does cardinality matter?
A label is indexed metadata such as job, host, or severity. Each unique combination of label values creates a separate stream, so putting something high-cardinality like a request ID or user ID in a label produces millions of streams and destroys performance. Keep labels to a small bounded set.
Do I need Promtail or can I use something else?
Promtail is the traditional agent and Grafana Alloy supersedes it, handling logs, metrics, and traces in one binary. Docker and Kubernetes can also ship directly through logging drivers, and systemd journals can be read natively rather than tailed from files.
Can Loki replace my existing log files?
It centralises them rather than replacing them. Local logs remain the authoritative copy and the thing you read when the network is down or Loki itself is broken. Shipping logs is about correlation across machines, not about deleting what is on disk.
How much storage does Loki need?
Far less than Elasticsearch for the same logs, frequently by an order of magnitude, because it compresses chunks and indexes almost nothing. Object storage such as S3 or MinIO is the usual backend, which makes retention a matter of lifecycle policy rather than disk provisioning.
Does Loki work with Grafana alerting?
Yes. LogQL supports metric queries that convert log lines into rates and counts, so you can alert on a rise in error lines exactly as you would on a Prometheus metric. That is one of the main reasons to run it alongside Prometheus rather than a separate log stack.