The awk Command Explained

The awk Command Explained

awk is a text-processing language, older than most of the tools that get compared to it, built around one central idea: split each line into fields, then let you write patterns and actions that operate on those fields. Where grep finds lines and sed transforms them as whole units, awk is the tool to reach for once you need to think in columns, whether that means extracting a specific field, summing a numeric column, or reformatting delimited data.

The pattern-action model

Every awk program is a series of pattern { action } pairs. For each line of input, awk checks whether the line matches the pattern, and if it does, runs the action:

awk '/error/ { print }' logfile.txt

This prints every line containing error, functioning similarly to grep error logfile.txt. The pattern is optional; if omitted, the action runs on every line. The action is also optional; if omitted, matching lines are printed by default, exactly like grep.

Fields and the print statement

awk automatically splits each input line into fields on whitespace by default, accessible as $1, $2, $3, and so on. $0 refers to the entire unmodified line.

echo "alice 25 engineer" | awk '{print $1}'
# Output: alice

echo "alice 25 engineer" | awk '{print $2, $3}'
# Output: 25 engineer

This is the single most common awk use case: pulling one or two columns out of structured output, such as the output of ps, df, or ls -l.

# Print just the memory usage column from free
free -h | awk 'NR==2 {print $3}'

# Print filenames and sizes from ls -l
ls -l | awk '{print $9, $5}'

Changing the field separator

The default field separator is whitespace, but many real files use commas, colons, or other delimiters. Set a custom separator with -F:

# CSV file
awk -F, '{print $2}' data.csv

# /etc/passwd is colon-delimited
awk -F: '{print $1}' /etc/passwd

GNU awk also accepts a regular expression as the field separator, useful when fields are separated by inconsistent whitespace or multiple possible delimiters:

awk -F'[,;]+' '{print $1}' data.txt   # split on commas or semicolons

Built-in variables

awk provides several variables that update automatically as it processes input, without any setup required:

VariableMeaning
$0The entire current line
$1, $2, …Individual fields on the current line
NFNumber of fields on the current line
NRNumber of the current line (record number) across the whole input
FSThe field separator (same as setting it with -F)
OFSThe output field separator, used when printing multiple fields

NF, the field count, is particularly useful for grabbing the last column regardless of how many fields a line has:

echo "one two three four" | awk '{print $NF}'
# Output: four

NR is commonly used to act on a specific line number or a range of lines:

awk 'NR==1' file.txt        # print only the first line (often a header)
awk 'NR>1' file.txt         # print every line except the first (skip a header)
awk 'NR>=5 && NR<=10' file.txt  # print lines 5 through 10

BEGIN and END blocks

BEGIN runs once before any input is read, and END runs once after all input has been processed. Both are useful for setup and for printing summary results:

awk 'BEGIN {print "Starting scan..."} {print $1} END {print "Done."}' file.txt

The most common use of END is accumulating a total across every line and printing it once at the end:

awk '{sum += $1} END {print "Total:", sum}' numbers.txt

sum does not need to be declared or initialized beforehand; awk treats a new numeric variable as zero the first time it is referenced.

Conditionals inside the action

awk supports if/else inside an action block, letting you filter and transform in the same pass:

awk '{ if ($3 > 100) print $1, "is over the limit" }' data.txt

Comparisons can also serve directly as the pattern, without an explicit if:

awk '$3 > 100 {print $1}' data.txt

This prints the first field of any line where the third field is greater than 100, a compact combination of pattern and action that is common in real awk one-liners.

Common real-world one-liners

# Count lines (same as wc -l)
awk 'END {print NR}' file.txt

# Print unique values in a column
awk '{print $1}' file.txt | sort -u

# Sum a column, e.g. total size from du -k
du -k * | awk '{sum += $1} END {print sum, "KB total"}'

# Print lines longer than 80 characters
awk 'length($0) > 80' file.txt

# Print the second-to-last field
awk '{print $(NF-1)}' file.txt

# Reformat colon-delimited output as comma-delimited
awk -F: 'BEGIN {OFS=","} {print $1, $3}' /etc/passwd

That last example sets OFS, the output field separator, so that fields printed with commas between them (via the comma in the print statement) come out comma-separated in the output, even though the input was split on colons.

awk vs sed vs grep

The three classic Unix text tools each have a natural home. grep answers “which lines match this pattern.” sed answers “how do I transform or delete matching lines.” awk answers “how do I extract, compute, or rearrange specific columns.” Real shell pipelines frequently chain them together, for example filtering with grep, cleaning formatting with sed, and extracting a final numeric value with awk, each tool doing the part it is best suited for rather than forcing one tool to do everything.

Frequently Asked Questions

What is awk used for?

awk is a text-processing language designed around splitting each line of input into fields (by default, on whitespace) and letting you write patterns and actions that operate on those fields. It is most useful when your data is organized in columns, such as command output, CSV-like files, or log lines, and you want to extract, filter, rearrange, or compute across specific columns rather than treat each line as an undivided block of text. Common uses include printing a specific column from command output, summing a numeric column, and reformatting delimited data.

How do I print a specific column with awk?

Use the field variable for the column you want: awk ‘{print $2}’ file.txt prints the second whitespace-separated field of every line. Fields are numbered starting at 1, and $0 refers to the entire line unmodified. To print multiple columns, list them separated by a comma inside the print statement, such as awk ‘{print $1, $3}’ file.txt, which prints the first and third fields separated by a space.

How do I change the field separator in awk?

Use the -F flag to set the field separator for input, for example awk -F, ‘{print $2}’ file.csv treats commas as the field delimiter instead of the default whitespace. The separator can be any single character or, in GNU awk, a regular expression, such as awk -F ’:+’ to split on one or more colons. This is essential for working with CSV files, /etc/passwd-style colon-delimited files, and any other structured text that is not whitespace-separated.

What is the difference between awk and sed?

sed operates on whole lines, primarily for substitution and deletion, and has no real concept of columns or fields. awk splits each line into fields and is built around computing and extracting values from specific columns, along with basic arithmetic and string operations. If your task is find-and-replace or deleting matching lines, sed is the simpler tool. If your task involves extracting a specific column, summing a numeric field, or reformatting delimited data, awk is a better fit. The two are often used together in a pipeline, with sed cleaning up formatting and awk extracting the final values.

How do I sum a column of numbers with awk?

Use a running total variable inside the main block, and print it in an END block that runs once after all lines are processed: awk ‘{sum += $1} END {print sum}’ file.txt adds up the first field of every line and prints the total. The sum variable does not need to be declared in advance; awk initializes numeric variables to zero automatically the first time they are used. This pattern generalizes to averages, counts, and other running calculations across a file.

What do NR and NF mean in awk?

NR is the number of records (lines) processed so far, effectively a running line counter, commonly used to print or act on a specific line number: awk ‘NR==5’ file.txt prints only line 5. NF is the number of fields on the current line, useful for printing the last column regardless of how many columns a line has, since $NF always refers to the last field: awk ‘{print $NF}’ file.txt. Both are built-in variables that update automatically as awk processes each line, with no setup required.