lnav: Log Analysis Without grep Gymnastics

lnav: Log Analysis Without grep Gymnastics

Most log investigation looks like this:

zcat /var/log/nginx/access.log.*.gz | grep " 500 " | awk '{print $1}' | sort | uniq -c | sort -rn

That works and it is a lot of typing, it loses the timestamps, and correlating it with the error log means doing the whole thing again and reading two outputs side by side.

lnav parses logs rather than treating them as text.

sudo apt install lnav
lnav /var/log/nginx/

What parsing buys you

It ships format definitions for syslog, nginx, Apache, and a few dozen others. Once a format is recognised:

Timestamps become real. It can sort, filter by time, and jump to a moment.

Fields become addressable. Status code, client address, request path, and response time are columns rather than whitespace-separated positions.

Several files merge into one timeline. This is the feature that matters most. Point it at a directory and the access log and the error log interleave in chronological order, so you see the 500 and the traceback that produced it adjacent to each other.

Compressed and rotated files are handled transparently. gzip, bzip2, and zstd, in the correct order.

lnav /var/log/nginx/         # everything, merged
lnav /var/log/syslog*        # including rotated archives
journalctl -f | lnav         # pipe the journal in
lnav -r /var/log/nginx/      # follow, like tail -f
e / E          next / previous error
w / W          next / previous warning
i              histogram view
/text          search
n / N          next / previous match
t              switch to text view
q              back
:goto 15:30    jump to a time

e is the one to learn first. It jumps to the next line the format definition classifies as an error, across every file, in time order. On a busy server that is considerably faster than searching for likely strings.

The histogram (i) shows message volume over time, broken down by severity. A spike is immediately visible, and you press Enter on it to jump there. Finding when something started going wrong usually takes seconds.

Filtering

:filter-in 500
:filter-out /healthz
:filter-out 127.0.0.1
:reset-session

Filters stack, and they are live rather than a new pipeline. :filter-out /healthz removes health check noise, which on a monitored service can be most of the file.

Press TAB to see active filters and remove individual ones.

SQL

The feature that distinguishes it.

Each recognised format exposes its fields as columns in a SQLite virtual table.

;SELECT * FROM access_log LIMIT 5;
;SELECT c_ip, count(*) AS n
 FROM access_log
 WHERE sc_status >= 500
 GROUP BY c_ip
 ORDER BY n DESC
 LIMIT 10;

That is the pipeline from the top of this article, written as the question you were actually asking.

-- slowest endpoints
;SELECT cs_uri_stem, count(*) AS hits,
        avg(time_taken) AS avg_ms,
        max(time_taken) AS max_ms
 FROM access_log
 GROUP BY cs_uri_stem
 HAVING hits > 100
 ORDER BY avg_ms DESC
 LIMIT 20;

-- error rate per hour
;SELECT strftime('%Y-%m-%d %H:00', log_time) AS hour,
        count(*) AS total,
        sum(sc_status >= 500) AS errors
 FROM access_log
 GROUP BY hour
 ORDER BY hour;

-- what a specific client did before failing
;SELECT log_time, cs_method, cs_uri_stem, sc_status
 FROM access_log
 WHERE c_ip = '203.0.113.42'
 ORDER BY log_time;

Column names vary by format. ;.schema lists them.

Results appear in a pane you can navigate, and selecting a row jumps to it in the log.

The journal

journalctl -f | lnav
journalctl -u nginx --since "2 hours ago" | lnav
journalctl -k -b -1 | lnav          # previous boot's kernel messages

Piping works everywhere and is the simplest approach. The journal itself already provides much of the filtering lnav adds to plain files, so the main gain is the histogram, the SQL, and merging journal output with plain text logs from the same incident.

Our journalctl query builder covers the journal side.

Custom formats

For an application log lnav does not recognise:

{
  "myapp_log": {
    "title": "MyApp log",
    "regex": {
      "std": {
        "pattern": "^(?<timestamp>\\d{4}-\\d{2}-\\d{2}T\\d{2}:\\d{2}:\\d{2}Z) (?<level>\\w+) \\[(?<component>[^\\]]+)\\] (?<body>.*)$"
      }
    },
    "level-field": "level",
    "level": {
      "error": "ERROR",
      "warning": "WARN",
      "info": "INFO",
      "debug": "DEBUG"
    },
    "value": {
      "component": { "kind": "string", "identifier": true }
    },
    "sample": [
      { "line": "2026-09-14T10:30:00Z ERROR [database] connection refused" }
    ]
  }
}
mkdir -p ~/.config/lnav/formats/installed
cp myapp.json ~/.config/lnav/formats/installed/
lnav -C     # validate the format definition

lnav -C checks your regex against the samples and reports what does not match, which saves a great deal of guessing.

Named capture groups become SQL columns, so component is immediately queryable.

The sample array is mandatory. lnav uses it to verify the format and refuses to load a definition whose samples do not match.

Where it fits

Good for: investigating an incident across several log files, finding when something started, answering aggregate questions without writing a pipeline, and reading logs on a server where you have nothing else installed.

Not for: long-term storage, alerting, dashboards, or logs from many hosts. That is Loki, Elasticsearch, or a hosted service. Our monitoring guide covers that layer.

lnav is a viewer. It reads what is on the disk in front of you, and for a single server or a single incident that is usually exactly the scope of the problem.

The practical case for installing it: the next time you are on a server at an awkward hour trying to work out what happened, lnav /var/log/nginx/ and pressing e gets you further in ten seconds than the pipeline you were about to write.

Frequently Asked Questions

What does lnav do that less and grep do not?

It understands log formats, so it parses timestamps and fields rather than treating each line as text. That lets it merge several files into one chronological timeline, filter by time, highlight errors automatically, and run SQL queries against the parsed fields.

Can lnav read compressed and rotated logs?

Yes. Point it at a directory or a glob and it handles gzip, bzip2, and zstd transparently, then merges everything into a single timeline in the correct order. This is considerably easier than decompressing rotated files manually to search across a date boundary.

How does SQL work on a log file?

lnav exposes the parsed fields of each recognised format as columns in a virtual SQLite table. You can then write ordinary SQL with WHERE, GROUP BY, and aggregate functions against them, which turns questions like top clients by error count into one query.

Does lnav work with systemd journal logs?

Yes, either by piping journalctl output into it or by using its journald support directly. Piping works everywhere and is the simplest approach, and the journal itself already provides much of the filtering that lnav adds to plain text files.

Can I use lnav on a log format it does not recognise?

Yes, by writing a format definition in JSON that specifies a regular expression to extract fields and which one is the timestamp. Place it in the formats directory under your lnav configuration and it applies automatically to matching files.

Is lnav suitable for very large log files?

It handles multi-gigabyte files, indexing incrementally so you can start reading before the whole file is processed. For repeated analysis across many gigabytes, loading into a real database or a log aggregation system is more appropriate, since lnav is a viewer rather than a store.