uniq Command Explained
uniq removes duplicate lines from text, but with one crucial limitation that trips up nearly everyone the first time: it only catches duplicates that are already sitting right next to each other.
Basic usage
cat colors.txt
# red
# red
# blue
# green
# green
# green
uniq colors.txt
# red
# blue
# green
Each run of consecutive identical lines collapses down to a single instance. Notice that this only worked here because the duplicate lines happened to already be grouped together in the original file.
Why uniq requires sorted input to work properly
cat unsorted_colors.txt
# red
# blue
# red
# green
uniq unsorted_colors.txt
# red
# blue
# red
# green
# (nothing was removed at all! the two "red" lines are not
# adjacent to each other, so uniq never notices they match)
This is the single most important thing to understand about uniq: it only ever compares each line to the one directly before it. If two identical lines are separated by anything else in between, uniq has no mechanism to detect that they are duplicates, since it never looks beyond the immediately preceding line.
sort unsorted_colors.txt | uniq
# blue
# green
# red
Sorting first guarantees every occurrence of an identical line ends up adjacent to every other occurrence of that same line, which is exactly why sort file.txt | uniq is by far the most common way these two commands appear together, and arguably why they are so often thought of as a single combined idiom rather than two separate tools.
Counting occurrences: uniq -c
sort access.log | uniq -c | sort -rn
# 342 192.168.1.10
# 128 192.168.1.55
# 12 10.0.0.3
This three-command pipeline is an extremely common pattern for quick frequency analysis: sort to group identical lines together, uniq -c to collapse each group into a count, and a second sort -rn to order the results from most frequent to least frequent. Here it is being used to find which IP address appears most often in a log file, but the same pattern works for counting occurrences of anything: error messages, usernames, HTTP status codes.
Showing only duplicated lines
sort names.txt | uniq -d
# only lines that appeared MORE than once, nothing that
# was unique in the original data
-d filters the output down to only lines that had at least one duplicate, which is useful when you specifically want to find repeated entries in a dataset without also seeing every unique entry mixed in.
Showing only lines with no duplicates
sort names.txt | uniq -u
# only lines that appeared EXACTLY once, the opposite of -d
-u is the mirror image of -d, isolating entries that are genuinely one-of-a-kind, with no duplicate anywhere else in the dataset.
Case-insensitive comparison
sort -f names.txt | uniq -i
-i makes uniq’s duplicate comparison case-insensitive, treating “Error” and “error” as the same line for deduplication purposes. Note that the sorting step should generally also account for case (sort -f) to ensure case-variant duplicates actually end up adjacent to each other in the first place before uniq ever sees them.
Skipping fields or characters before comparing
uniq -f 1 data.txt
# ignores the first whitespace-separated field on each line
# when deciding whether two lines are duplicates
uniq -s 5 data.txt
# ignores the first 5 characters of each line when comparing
-f N and -s N are useful when lines contain a varying prefix, such as a timestamp or sequence number, that should be ignored when determining whether the rest of the line’s content is actually a duplicate of another line’s content.
# Example: log lines with timestamps that make otherwise
# identical messages look unique to a naive comparison
# 10:15:01 Connection reset
# 10:15:04 Connection reset
uniq -s 9 log.txt
# skips the first 9 characters (the timestamp) before
# comparing, correctly identifying these as duplicates
Frequently Asked Questions
What does the uniq command do?
uniq filters out repeated lines from text, collapsing consecutive identical lines down to a single instance. It reads from a file or piped input and outputs each line once, but only removes duplicates that are directly adjacent to each other in the input.
Why does uniq only work correctly on sorted input?
uniq only compares each line to the one immediately before it, removing a duplicate only if it directly follows an identical line. If two identical lines exist in the input but are separated by other different lines, uniq has no way to detect that they are duplicates of each other, since it never compares non-adjacent lines. Sorting the input first with sort guarantees that every occurrence of an identical line ends up next to every other occurrence, which is why sort file.txt | uniq is by far the most common way uniq actually gets used in practice.
How do I count how many times each line appears using uniq?
Add the -c flag, such as sort file.txt | uniq -c, which prefixes each output line with the number of consecutive times it appeared in the input. This is a very common pattern for quick frequency analysis, such as counting how many times each unique IP address appears in a log file.
How do I show only the lines that appear more than once?
Use the -d flag (duplicates only), such as sort file.txt | uniq -d, which prints only lines that had at least one adjacent duplicate, omitting any line that appeared just a single time. This is useful for specifically identifying repeated entries without needing to scan through the full deduplicated output.
How do I show only the lines that appear exactly once, with no duplicates at all?
Use the -u flag (unique only), such as sort file.txt | uniq -u, which prints only lines that had no adjacent duplicates whatsoever, the opposite of -d. This isolates entries that are genuinely one-of-a-kind in the dataset.
Can uniq ignore case or ignore certain characters when comparing lines for duplicates?
Yes, -i makes the comparison case-insensitive, so “Error” and “error” are treated as the same line for deduplication purposes. -f N skips the first N whitespace-separated fields before comparing, and -s N skips the first N characters, both useful when you want to compare only part of each line rather than the entire line from the very beginning.