grep -rl "text-processing" ./blog

Text Processing on Linux

grep, sed, awk, cut, and tr: the small tools that compose into pipelines capable of reshaping almost any text stream.

10 articles

fzf and ripgrep Explained: Faster Searching in the Terminal ripgrep searches file contents faster than grep and respects .gitignore; fzf turns any list into an interactive fuzzy finder. How each works, the shell keybindings that change daily terminal use, and how to combine them. nl, rev, and fold: Three Small Text Utilities Worth Knowing Numbering lines with control over the format, reversing characters on a line, and wrapping long lines to a width. Each solves one problem that is awkward otherwise. comm Command Explained: Comparing Two Sorted Files comm tells you what is only in the first file, only in the second, and in both. It is faster and clearer than diff for set comparisons, with one strict requirement. paste Command Explained: Merging Files Side by Side paste joins lines from several files into columns, or folds a single column into rows. It is the counterpart to cut and the quickest way to rebuild a table from separate pieces. column Command Explained: Making Terminal Output Readable Ragged output from cut, awk, or a CSV becomes an aligned table with one extra command. column is the least known and most immediately satisfying formatting tool on Linux. wc Command Explained: Counting Lines, Words, and Bytes wc counts what passes through it. Most of its real use is one flag, -l, at the end of a pipeline, and there are a few traps worth knowing about what it counts as a line. The cut Command Explained: Extracting Columns from Text cut slices lines of text by delimiter, field number, or character position. It is the fastest way to pull one column out of a CSV, a log line, or /etc/passwd without reaching for awk. The tr Command Explained: Translating and Deleting Characters tr transforms characters in a stream: uppercasing text, squeezing repeated spaces, deleting carriage returns, and converting delimiters. Small tool, constant usefulness in pipelines. The awk Command Explained awk is a field-based text processing language built into every Linux system, ideal for extracting columns, summing values, and reformatting structured text. This guide covers its pattern-action model, built-in variables, and common one-liners. The sed Command Explained sed is the standard Linux stream editor for find-and-replace, line deletion, and text transformation from the command line. This guide covers substitution syntax, in-place editing, addressing specific lines, and common real-world patterns.