Wildcards vs Regular Expressions Explained

Wildcards vs Regular Expressions Explained

Shell wildcards and regular expressions are two different pattern-matching systems that happen to reuse some of the same characters, * most notably, which is exactly why they get confused so often. Understanding where each one applies clears up a lot of “why isn’t this pattern matching what I expected” confusion.

Two different systems, one confusing overlap

Shell wildcards, also called globbing, are a simple pattern language built into the shell itself, used mainly for matching filenames on the command line. Regular expressions (regex) are a separate, more powerful pattern language used by text-processing tools like grep, sed, and awk to match content within text. Both use * as a special character, but it means something structurally different in each.

Shell wildcards (globbing)

ls *.txt          # any filename ending in .txt
ls file?.txt       # file1.txt, file2.txt, fileA.txt, but not file10.txt (exactly one character)
ls file[123].txt   # file1.txt, file2.txt, file3.txt specifically
ls file[0-9].txt   # any single digit in that position

Three special characters cover almost all globbing: * matches any sequence of characters, including zero characters. ? matches exactly one character, no more, no less. [ ] matches any single character from the set or range listed inside, such as [abc] or [0-9].

A crucial detail: the shell expands these patterns itself, before the command being run ever sees them. ls *.txt does not pass the literal text *.txt to ls, it passes an already-expanded list of every matching filename. This is why globbing only works for command-line arguments the shell processes, and has no effect at all on text inside a file’s contents.

echo *.txt
# expands to: notes.txt readme.txt todo.txt   (whatever actually matches)

Regular expressions

Regular expressions are a different, more expressive system, used by tools that search or match patterns within text content, not filenames on the command line:

grep "^Error" logfile.txt        # lines starting with "Error"
grep "[0-9]\{3\}" logfile.txt     # lines containing three consecutive digits
grep "a.*b" file.txt              # lines containing "a", then anything, then "b"

Regex supports concepts wildcards have no equivalent for at all: anchoring to the start (^) or end ($) of a line, repetition counts ({3} for exactly three, + for one or more), and alternation between multiple patterns (cat|dog matches either word).

Why * means something different in each

This is the core source of confusion. In a shell wildcard, * means “any sequence of characters.” In a regular expression, * means “zero or more of the immediately preceding character,” which is a fundamentally different rule:

# Wildcard: matches any file starting with "log"
ls log*

# Regex: a* means zero or more "a" characters, NOT "anything"
grep "a*b" file.txt
# matches: b, ab, aab, aaab...
# does NOT mean "anything followed by b"

To express “any sequence of characters” in regex, the equivalent is .*, not * alone: . matches any single character, and * means zero or more of it, so together they mean “any sequence of characters,” the regex way of expressing what a bare * means in a wildcard.

grep "a.*b" file.txt
# matches "a", then any characters (including none), then "b"
# this is the regex equivalent of a wildcard-style a*b

Where each one applies

Tool/contextUses
ls, cp, mv, rm (filename arguments)Wildcards
grep, sed, awk (matching text content)Regular expressions
find -nameWildcards
find -regexRegular expressions
Bash case statementsWildcards
[[ $var =~ pattern ]] in BashRegular expressions

find is a good example of both existing side by side: -name "*.txt" uses wildcard syntax, while -regex ".*\.txt" uses full regex syntax for the same general purpose, and mixing up which flag expects which syntax is a common find mistake.

find . -name "*.txt"              # wildcard: matches files ending in .txt
find . -regex ".*\.txt"            # regex: matches full path against the pattern

Basic vs extended regular expressions

Regex itself has variants. Basic regular expressions (BRE), the default for grep and sed, require a backslash before characters like +, ?, |, and ( ) for them to act as special regex syntax; without the backslash, they are literal characters. Extended regular expressions (ERE), enabled with -E, treat those characters as special by default, closer to what most people expect from regex in other languages:

grep 'a\+b' file.txt        # BRE: \+ means "one or more a"
grep -E 'a+b' file.txt       # ERE: + means "one or more a" without escaping

A practical rule of thumb

If you are typing a pattern directly on the command line to match filenames, for tools like ls, cp, or find -name, you are almost certainly using a wildcard, and only *, ?, and [ ] are meaningful. If you are matching or searching within the contents of a file, for tools like grep, sed, or awk, you are using regex, and a much larger set of characters (., *, +, ^, $, |, and more) carry special meaning. When a pattern that “should” match is not working, checking which system the tool you are using actually expects is often the fastest way to find the bug.

Frequently Asked Questions

What is the actual difference between wildcards and regular expressions?

Wildcards (also called globbing patterns) are a simple pattern-matching system used by the shell itself, mainly for matching filenames, with only a handful of special characters: *, ?, and [ ]. Regular expressions are a much more powerful and expressive pattern-matching language used by tools like grep, sed, and awk to match text within file contents, supporting things wildcards cannot do at all, such as anchoring to the start or end of a line, repetition counts, and alternation between multiple patterns. They look similar because both reuse some of the same characters, but * means something different in each system, which is precisely why confusing the two produces patterns that silently do not do what you expect.

Why does * mean something different in ls .txt versus grep “ab”?

In a shell wildcard like ls .txt, * means “any sequence of characters, including none,” matching filenames directly. In a regular expression like grep “ab”, * has a completely different meaning inherited from regex syntax: it means “zero or more of the immediately preceding character,” so a*b matches b, ab, aab, aaab, and so on, but does not mean “any characters” the way it would in a wildcard. This is the single most common source of confusion between the two systems, since the same character means something structurally different depending on which system is interpreting it.

Does grep use wildcards or regular expressions?

grep uses regular expressions, not shell wildcards, despite grep commonly being run alongside wildcard-based filename arguments in the same command line, which is part of why the two get conflated. grep “a.b” file.txt uses regex syntax: . matches any single character, and * means zero or more of the preceding character, so . together means “any sequence of characters,” the regex equivalent of what a plain * means in a wildcard. If you pass a plain * on its own to grep expecting wildcard behavior, it will not do what a filename wildcard would, since grep is interpreting it as regex syntax the whole time.

What do the shell wildcard characters actually mean?

Shell globbing supports three main patterns: * matches any sequence of characters (including zero characters), ? matches exactly one character, and [ ] matches any single character from the set or range listed inside the brackets, such as [abc] matching a, b, or c, or [0-9] matching any single digit. These are expanded by the shell itself before the command even runs; ls *.txt does not pass the literal string *.txt to ls, it passes the shell’s already-expanded list of matching filenames, which is why globbing only works for arguments the shell processes, not for text inside a file.

When would I need a regular expression instead of a wildcard?

Use a wildcard when matching filenames directly on the command line: ls, cp, mv, and rm all accept and expand wildcard patterns naturally through the shell. Use a regular expression when searching or matching patterns inside file contents, which is what grep, sed, awk, and find -regex are built around, since wildcards have no mechanism at all for expressing things like “a line starting with a digit” or “one or more repetitions of a specific character,” both of which are straightforward with regex syntax.

What is the difference between basic and extended regular expressions?

Basic regular expressions (BRE), the default in grep and sed, require a backslash before characters like +, ?, |, and ( ) for them to have their special meaning; without the backslash they are treated as literal characters. Extended regular expressions (ERE), enabled with grep -E or sed -E, treat those same characters as special by default, without needing a backslash, which is closer to the regex syntax most people are already familiar with from other programming languages. Practically, this means a pattern like a+b works differently depending on the mode: as BRE it matches the literal text a+b, and as ERE it matches one or more a characters followed by b.