gzip Explained
gzip is the compression tool most Linux users interact with without thinking about it, whether through .tar.gz archives, compressed log files, or the Content-Encoding: gzip header used across most of the web. It prioritizes a good balance of speed and compression ratio over squeezing out the smallest possible file.
Basic Usage
gzip largefile.log
This compresses largefile.log into largefile.log.gz and deletes the original file by default. To keep the original alongside the compressed version:
gzip -k largefile.log
Decompressing
gunzip largefile.log.gz
# equivalent to:
gzip -d largefile.log.gz
Both restore the original file and remove the .gz version (unless -k is used).
To view a compressed text file’s contents without fully decompressing it to disk:
zcat largefile.log.gz
zcat largefile.log.gz | grep "ERROR"
zcat decompresses to standard output, making it easy to pipe compressed logs directly into grep, less, or other text tools without an intermediate decompressed file.
Compression Levels
gzip -9 largefile.log # maximum compression, slowest
gzip -1 largefile.log # fastest, least compression
gzip accepts levels from -1 (fastest, largest output) to -9 (slowest, smallest output), with -6 as the implicit default. The difference between levels is a tradeoff between CPU time spent searching for compression matches and the resulting file size. For most everyday use the default is fine; reach for -9 when squeezing out extra space is worth the additional compression time, such as archiving data you won’t need to compress again.
gzip and tar
gzip only compresses a single data stream; it has no concept of directories or multiple files. tar solves the bundling problem separately, combining an entire directory tree into one archive stream, which gzip then compresses as a single unit:
tar -czvf archive.tar.gz /data
This is why compressed backups of directories are .tar.gz files rather than gzip being pointed at a folder directly (which isn’t possible; gzip operates on individual files).
gzip vs. bzip2 vs. xz
| Tool | Speed | Ratio | Typical Use |
|---|---|---|---|
gzip | Fastest | Moderate | Default choice, logs, HTTP compression |
bzip2 | Slower | Better | Middle ground, less common today |
xz | Slowest | Best | Software releases, long-term archives |
gzip’s speed advantage matters most in contexts where compression happens frequently or in real time, such as web servers compressing responses on the fly, or log rotation compressing files as part of a routine, frequent task. Where file size matters more than compression speed, and the operation happens rarely (like packaging a software release), xz is usually the better choice despite being considerably slower.
Practical Examples
# compress all .log files in a directory, keeping originals
gzip -k *.log
# decompress everything in a directory
gunzip *.gz
# check compression ratio achieved
gzip -l archive.gz
gzip -l prints the compressed size, uncompressed size, and ratio for a .gz file without decompressing it, useful for quickly checking how effective compression was on a particular file.
Frequently Asked Questions
What algorithm does gzip use for compression?
gzip uses the DEFLATE algorithm, which combines LZ77 dictionary-based compression with Huffman coding. DEFLATE is designed to be fast to compress and decompress while still achieving a reasonable compression ratio, which is why gzip remains a common default even though newer algorithms achieve smaller output sizes at the cost of speed.
Does gzip replace the original file after compressing it?
By default, yes. Running gzip file.txt produces file.txt.gz and deletes the original file.txt. To keep the original file alongside the compressed version, use the -k (keep) flag: gzip -k file.txt. This behavior differs from zip, which always produces a separate archive without touching the source files.
What do the gzip compression level flags like -1 and -9 mean?
gzip accepts a compression level from -1 (fastest, largest output) to -9 (slowest, smallest output), with -6 as the default balance point. Higher levels spend more CPU time searching for better compression matches, yielding smaller files at the cost of speed. For most everyday compression, the default level is a reasonable choice; -9 is worth using when file size matters more than the time spent compressing.
How do I decompress a .gz file?
Run gunzip file.txt.gz, or equivalently gzip -d file.txt.gz. Both restore the original uncompressed file and remove the .gz file, unless -k is used to keep the compressed copy. To view the contents of a compressed text file without fully decompressing it to disk, zcat file.txt.gz pipes the decompressed content directly to standard output.
Why do tar archives use gzip instead of gzip compressing the whole directory directly?
gzip only compresses a single file stream; it has no concept of directories, multiple files, or metadata like permissions and ownership. tar solves the bundling problem by combining a whole directory tree into one archive stream, which gzip then compresses as a single unit. This is why compressed directory backups are typically .tar.gz rather than gzip being applied directly to a folder.
How does gzip compare to bzip2 and xz in practice?
gzip is the fastest of the three at both compressing and decompressing, but produces the largest output files. bzip2 achieves better compression than gzip at a noticeably slower speed. xz achieves the best compression ratio of the three but is the slowest to compress. gzip remains the common default for everyday use and network protocols where speed matters; xz is preferred when minimizing file size matters more, such as software distribution archives.