bzip2 Explained
bzip2 occupies the middle ground in Linux compression: better ratios than gzip, slower than gzip, and generally a step behind xz on both fronts today. It’s still worth understanding since .tar.bz2 archives remain common enough to encounter regularly, even if it’s less often the first choice for new projects.
Basic Usage
bzip2 archive.tar
Compresses archive.tar into archive.tar.bz2 and removes the original by default. To keep the source file:
bzip2 -k archive.tar
Decompressing
bunzip2 archive.tar.bz2
# equivalent to:
bzip2 -d archive.tar.bz2
To view compressed content without writing a decompressed file to disk:
bzcat archive.txt.bz2 | less
How bzip2 Works, Briefly
bzip2 uses the Burrows-Wheeler transform (BWT), a block-sorting algorithm that rearranges data to group similar byte sequences together, making the result much more compressible by the move-to-front encoding and Huffman coding steps that follow. This is a fundamentally different approach from gzip’s dictionary-based DEFLATE algorithm, and it’s part of why the two tools have different speed and ratio characteristics.
Because BWT operates on fixed-size blocks rather than the whole file at once, bzip2 processes data in chunks, controlled by the block size flag.
Block Size
bzip2 -9 archive.tar # largest blocks (900 KB), default, best ratio
bzip2 -1 archive.tar # smallest blocks (100 KB), less memory, slightly worse ratio
Larger blocks generally compress slightly better since there’s more data available for the algorithm to find patterns in, at the cost of more memory during both compression and decompression. -9 is the default and rarely needs changing except on memory-constrained systems.
bzip2 with tar
tar -cjvf archive.tar.bz2 /data
tar -xjvf archive.tar.bz2
The -j flag tells tar to pipe the archive stream through bzip2. This is the standard way bzip2 gets used in practice, since it (like gzip and xz) only compresses a single data stream and has no concept of directories on its own.
bzip2 vs. gzip vs. xz
| Tool | Speed | Ratio | Notes |
|---|---|---|---|
gzip | Fastest | Moderate | Best when speed matters most |
bzip2 | Slower | Better than gzip | Middle ground, block-based |
xz | Slowest | Best | Best when size matters most |
The practical reality today is that most projects choosing a compressor for something other than raw speed lean toward xz, since it typically achieves better ratios than bzip2 at comparable or better speed for large files, particularly with multi-threaded compression via -T0. bzip2 still shows up in established pipelines and some tooling defaults, and remains fully supported, but it’s less often the deliberate choice for new setups.
When bzip2 Still Makes Sense
- Working with existing
.tar.bz2archives or pipelines already built around it - Environments where
xzisn’t available but a better ratio thangzipis needed - Some specific data types where BWT-based compression happens to outperform both alternatives
Frequently Asked Questions
What makes bzip2 different from gzip?
bzip2 uses the Burrows-Wheeler transform combined with move-to-front encoding and Huffman coding, a fundamentally different approach from gzip DEFLATE algorithm. This generally lets bzip2 achieve better compression ratios than gzip, particularly on larger text-like files, at the cost of being noticeably slower to both compress and decompress.
Does bzip2 replace the original file like gzip does?
Yes, by default bzip2 file.txt produces file.txt.bz2 and removes the original file.txt. Use bzip2 -k file.txt to keep the original file alongside the compressed version, matching the same -k convention used by gzip.
How do I decompress a .bz2 file?
Run bunzip2 file.txt.bz2, or equivalently bzip2 -d file.txt.bz2. To view the contents of a compressed text file without writing a decompressed copy to disk, bzcat file.txt.bz2 streams the decompressed content directly to standard output.
Is bzip2 still commonly used today?
bzip2 remains supported and available on virtually every Linux system, but its use has declined somewhat as xz has become the preferred choice for high-ratio compression, offering better ratios than bzip2 at broadly comparable or sometimes faster speeds for large files. bzip2 still appears in some established tooling and legacy pipelines, and some file types compress better under bzip2 specific block-based approach, but for new projects xz is usually the more common recommendation when gzip speed is not fast enough.
What does the block size flag in bzip2 control?
bzip2 processes data in blocks, with the block size configurable from -1 (100 KB blocks) to -9 (900 KB blocks, the default). Larger blocks generally improve compression ratio slightly at the cost of more memory usage during compression and decompression, since bzip2 works within one block at a time rather than across the whole file.
How does bzip2 compression ratio and speed compare to gzip and xz in practice?
On typical text data, bzip2 usually compresses noticeably smaller than gzip but takes longer to do so. Compared to xz, bzip2 is often faster but produces slightly larger output, though the exact difference varies by data type. For most current use cases, the practical choice comes down to gzip for speed, xz for the smallest files, with bzip2 occupying a middle ground that is less commonly chosen for new projects specifically but remains a reasonable option.