split Command Explained: Breaking Large Files Into Pieces
split breaks a file into smaller pieces. cat puts them back. That is the whole tool, and it solves a specific set of problems cleanly.
Splitting
# By size
split -b 100M bigfile.iso part_
# By line count
split -l 10000 huge.log chunk_
# Into a fixed number of pieces
split -n 4 bigfile.dat piece_
The last argument is the filename prefix. Without it you get files called xaa, xab, and so on.
split -b 100M -d bigfile.iso part_
# part_00 part_01 part_02
-d gives numeric suffixes, which sort more predictably and read better than the default aa, ab, ac.
Reassembling
cat part_* > bigfile.iso
The glob works because split names pieces so that lexical order matches the original order. That is the reason for the naming scheme.
Always verify afterwards:
sha256sum bigfile.iso
A truncated transfer of one piece produces a reassembled file that is the wrong size and gives no other warning.
Where it is genuinely useful
Transfer size limits. Email attachments, upload caps, or filesystems with a maximum file size.
split -b 1900M archive.tar.gz archive.tar.gz.part_
Parallel processing. Split a large input and process the pieces concurrently:
split -n l/8 -d huge.csv chunk_
ls chunk_* | xargs -P 8 -I{} ./process.sh {}
Note -n l/8 rather than -n 8. The l/ prefix means split into 8 pieces without breaking lines, which is what you want for text. Plain -n 8 splits by byte count and will cut a line in half.
Log handling. Breaking an enormous log into pieces small enough for a tool that struggles with the whole thing.
Bytes versus lines
The distinction matters more than it first appears.
-b splits at an exact byte offset with no regard for content. A line spanning the boundary is cut in two, so neither piece is valid text on its own.
-l splits at line boundaries, so each piece is valid text.
For anything you intend to process as text, use -l or -n l/N.
The CSV caveat
Splitting a CSV by lines gives you pieces where only the first has a header:
# Keep the header, split the rest, re-add the header to each piece
head -1 data.csv > header.txt
tail -n +2 data.csv | split -l 100000 -d - chunk_
for f in chunk_*; do cat header.txt "$f" > "$f.csv" && rm "$f"; done
split has no concept of a header, so this is on you.
Piped input
tar czf - /var/log | split -b 100M -d - backup.tar.gz.part_
The - reads standard input. Reassemble and extract:
cat backup.tar.gz.part_* | tar xzf -
Our tar builder covers the archive side, and note that the pieces are not individually extractable.
Frequently Asked Questions
How do I reassemble a split file?
Concatenate the pieces in order with cat, redirecting to the output file. Because split names pieces so they sort correctly, a shell glob puts them back in the right sequence automatically.
What is the difference between -b and -l?
The -b flag splits by size in bytes, which can cut a line in half. The -l flag splits by line count, keeping lines intact. Use -l for text you intend to process and -b for binary or transfer size limits.
How do I keep the suffixes numeric instead of alphabetic?
Pass -d for numeric suffixes, giving names ending 00, 01, 02 rather than aa, ab, ac. Add -a to control the suffix length when you expect more pieces than two characters allow.
Can split handle piped input?
Yes, it reads standard input when no file is given, which is how you split the output of a command without writing an intermediate file. This is common for compressing and splitting in one pass.
Does splitting a compressed file work?
Splitting works on any file, and the individual pieces are not independently decompressible. You must reassemble the whole thing before decompressing, so treat the pieces as fragments rather than as usable files.
What if I need each piece to stay valid on its own?
Split by lines with -l on a text format where each line is self-contained, or use a format-aware tool. For CSV you also need to repeat the header into each piece, which split does not do for you.