Basic and Extended Regex

grep: Searching Text

Chapter 3 ยท Basic and Extended Regex

So far the patterns were plain words. A pattern can also describe text: “a number”, “a line that starts with a date”, “the same word twice”. That description language is the regular expression, and grep has two spellings of it. This chapter covers grep's own two spellings and where they trip people up. For regular expressions as a subject in their own right (and other tools that use them), see Learning Regular Expressions.

Run on two versions of grep
Every example was run on GNU grep 3.0 (Git Bash) and GNU grep 3.11 (WSL), on the same practice file. Output was identical except for three examples, which come up where they matter: how a Windows line ending (CRLF) interacts with $, and a warning about * at the start of a pattern. Anything marked GNU works in GNU grep (Linux, WSL, Git Bash) and may not in macOS/BSD grep, which I did not try.

One Language, Two Spellings

By default grep reads the pattern as a basic regular expression (BRE). With -E it reads an extended one (ERE). They describe the same things; they differ in which punctuation characters are operators and which are plain text. Learn the shared core first, then the difference.

Practice File

Paste this into an empty folder. It makes lines.txt, 20 lines chosen to trip up each idea in this chapter.

cat > lines.txt <<'EOF' cat concatenate The cat sat on the mat cot cut cat c.t color colour a+b=c aaa abab abab price: $19.99 call 555-0142 now 2026-10-08 ERROR disk full 2026-10-08 WARN low memory the the end [note] keep this path C:\temp\file hello world indented line file.txt and fileXtxt end. EOF

Where: Anchors

^ matches the start of a line and $ the end. They match a position, not a character.

grep '^the' lines.txt
the the end

Only the line that starts with a lowercase the: “The cat sat” has a capital T, and other lines have the in the middle.

grep 'end\.$' lines.txt
end.

Together, ^$ is an empty line, the classic way to count or remove blank lines:

grep -c '^$' lines.txt
1

What: One Character and Repeats

. matches any one character. -o prints just what matched, which is the best way to see what a pattern really did:

grep -o 'c.t' lines.txt
cat cat cat cot cut cat c.t

The last match is a literal c.t (the dot matched a dot). Put * after something to mean “zero or more of it”:

grep -E '*a' lines.txt | head -2; echo status $?
grep: warning: * at start of expression cat concatenate status 0

Two traps with *. First, it is greedy: it takes as much as it can. .* (“anything”) between two letters swallows everything up to the last match on the line:

grep -o 'c.*t' lines.txt
cat concatenat cat sat on the mat cot cut cat c.t

Look at the second line of output: it starts at the first c in concatenate and runs to the last t in the word. Second, “zero or more” includes zero, so a pattern made only of optional things matches every line, even ones with no x at all:

grep -c 'x*' lines.txt; wc -l < lines.txt
20 20

All 20 lines matched x*. If you mean “one or more”, say so (xx*, or x\+ / x+ below).

Sets: Bracket Expressions

Square brackets match one character from a set. A ^ straight after the [ reverses it, and a dash makes a range:

grep -o 'c[aou]t' lines.txt
cat cat cat cot cut cat
grep -o 'c[^aou]t' lines.txt
c.t
grep -o '[0-9][0-9]*' lines.txt
19 99 555 0142 2026 10 08 2026 10 08

Inside brackets almost everything stops being special, so to match a literal dot you can write [.] instead of \.. A ] must come first in the set, and a backslash is just a backslash:

grep '[.]' lines.txt
cot cut cat c.t price: $19.99 file.txt and fileXtxt end.
grep '[]]' lines.txt
[note] keep this
grep '[\]' lines.txt
path C:\temp\file

Named classes

Ready-made sets, written with double brackets, which are defined by the locale (the language settings), so they follow that language's idea of a letter or a digit; a range such as [a-z] depends on how the locale orders letters:

ClassMatches
[[:digit:]]0-9
[[:alpha:]] / [[:alnum:]]letters / letters and digits
[[:upper:]] / [[:lower:]]upper / lower case letters
[[:space:]]space, tab, newline and similar
[[:punct:]]punctuation
grep -o '[[:digit:]]\+' lines.txt
19 99 555 0142 2026 10 08 2026 10 08

The class name sits inside a bracket expression, so the brackets are doubled. Writing it with one pair is a very common slip, and grep says so:

grep '[:digit:]' lines.txt; echo status $?
grep: character class syntax is [[:space:]], not [:space:] status 2

The Difference: Basic vs Extended

Here is the whole difference. In a basic expression the characters + ? | ( ) { } are ordinary text, and you put a backslash in front of them to make them operators. In an extended expression (-E) it is the other way round: they are operators, and the backslash makes them ordinary text.

You wantBasic (default)Extended (-E)
one or morea\+a+
optionala\?a?
either / ora\|ba|b
group\(ab\)(ab)
between 2 and 4a\{2,4\}a{2,4}
a literal plus signa+a\+

(The backslash forms in the basic column are GNU extensions: they work in GNU grep but not in every grep. Standard basic expressions have no + ? | at all, which is a reason to prefer -E in scripts that may run elsewhere.) Every one of these has a version that silently does the wrong thing. “Optional u”, first as basic written like extended, which looks for a real question mark:

grep 'colou?r' lines.txt; echo status $?
status 1
grep 'colou\?r' lines.txt
color colour
grep -E 'colou?r' lines.txt
color colour

Status 1, no error, no output: grep was happily looking for the text colou?r. And the mirror image, an extended pattern with a backslash, which is also a literal:

grep -E 'colou\?r' lines.txt; echo status $?
status 1

The plus sign is the clearest case, because one line of our file really contains a+b:

grep 'a+b' lines.txt
a+b=c
grep -E 'a+b' lines.txt
abab abab
grep -E 'a\+b' lines.txt
a+b=c

The first found the literal text a+b. The second found the line with abab in it, because a+b in extended means “one or more a, then b”. The third asked for a literal plus again. The same trap with either/or:

grep 'ERROR|WARN' lines.txt; echo status $?
status 1
grep 'ERROR\|WARN' lines.txt
2026-10-08 ERROR disk full 2026-10-08 WARN low memory
grep -E 'ERROR|WARN' lines.txt
2026-10-08 ERROR disk full 2026-10-08 WARN low memory
The rule I would give a new user
Use -E whenever the pattern needs + ? | ( ) { }, and then write them without backslashes. Reading one set of rules is easier than two. Reserve the basic form for simple patterns with no operators.

Groups, Repeats of Groups and Counts

Brackets repeat one character. A group repeats a whole piece, and {n} or {n,m} says how many times:

grep -E '(ab){2}' lines.txt
abab abab
grep '\(ab\)\{2\}' lines.txt
abab abab
grep -E '[0-9]{3}-[0-9]{4}' lines.txt
call 555-0142 now
grep '[0-9]\{3\}-[0-9]\{4\}' lines.txt
call 555-0142 now

Those are the same searches written both ways. The extended ones are shorter and easier to read.

Back-references: Matching Something Again

A group remembers what it matched. \1 means “the same text as group 1 matched”, \2 group 2, and so on. It works in both forms (in -E it is a GNU extension; standard extended expressions do not have it):

grep -E '(ab)\1' lines.txt
abab abab
grep '\(ab\)\1' lines.txt
abab abab

That found abab: the group matched ab and \1 demanded ab again. The classic use is finding a repeated word:

grep -oE '\b([a-z]+) \1\b' lines.txt
abab abab the the

\b is a word boundary (below), [a-z]+ is a word, then a space, then \1, the same word again. It found the the, and also abab abab, whose first abab is a word repeated.

Word Edges and Shorthands (GNU)

WriteMeaning
\ba word boundary: between a word character and a non-word character, or the line edge
\< and \>the start and the end of a word
\BNOT a word boundary (inside a word)
\w / \Wa word character (letter, digit, underscore) / anything else
\s / \Swhite space / anything else
grep -o '\bcat\b' lines.txt
cat cat cat
grep -o '\<cat\>' lines.txt
cat cat cat

Both found the whole word cat three times (not the one inside concatenate). \B is the reverse, the one inside a word:

grep -o '\Bcat' lines.txt
cat

\w+ is a neater way to say “a word” than [a-zA-Z0-9_]+:

grep -o '\w\+' lines.txt | head -3
cat concatenate The

Note that \d is not “a digit” in grep's basic or extended expressions. It is only a digit in Perl mode (-P, Chapter 5). In basic and extended it quietly means the letter d; use [0-9] or [[:digit:]].

Matching a Special Character Literally

Put a backslash in front of a character that has a meaning to make it plain, or use -F (Chapter 1) when the whole pattern is plain text. Unescaped, the dot in a file name matches anything:

grep -o 'file.txt' lines.txt
file.txt fileXtxt
grep -o 'file\.txt' lines.txt
file.txt

The dollar sign needs particular care because $ is an anchor in extended expressions wherever it appears, but in a basic expression it is an anchor only at the very end of the pattern, and plain text anywhere else:

grep '$19' lines.txt; echo status $?
price: $19.99 status 0
grep -E '$19' lines.txt; echo status $?
status 1
grep -E '\$19' lines.txt
price: $19.99

The basic one found the price by accident of that rule; the extended one looked for “the end of a line, then 19”, which can never match. Escape it (\$) and both agree.

Which Match Wins: Longest

When a pattern could match in more than one way, grep takes the longest match at the leftmost position, whatever order you listed the alternatives in:

echo foobar | grep -oE 'foo|foobar'
foobar
echo foobar | grep -oE 'foobar|foo'
foobar

Perl-style engines (Chapter 5) take the first alternative that works instead, so the same patterns give different answers there. It matters whenever you use -o or edit with the match.

Windows Line Endings

A file written on Windows ends each line with two characters, carriage return and line feed (CRLF). On Linux, grep sees the carriage return as part of the line, so $ (the end of the line) comes after it, and done$ does not match done:

printf 'done\r\n' | grep 'done$'; echo status $?
status 1
printf 'done\r\n' | grep $'done\r$' | cat -A
done^M$

That is WSL's grep 3.11, behaving as it would on any Linux server. Git Bash's grep 3.0 behaves differently: it is a Windows build and evidently drops the carriage return, so the same first command printed done and exited with status 0 there, and the second printed nothing. So a pattern that works in Git Bash can fail when the same file reaches a Linux machine. If a $ pattern mysteriously fails on a file from Windows, look for the carriage returns (cat -A file shows them as ^M), and either match them (as the second command did, using bash's $'...\r' quoting) or convert the file first.

When grep Complains

A pattern that is not valid gives a message and status 2:

grep -E '(' lines.txt; echo status $?
grep: Unmatched ( or \( status 2
grep '[' lines.txt; echo status $?
grep: Invalid regular expression status 2

In a basic expression an unmatched ( is only a bracket character, so there is nothing to complain about (status 1, nothing found):

grep '(' lines.txt; echo status $?
status 1

Newer GNU grep also warns about patterns that are legal but usually a mistake. A * at the start of an extended pattern has nothing in front of it to repeat. In 3.11 (shown here) grep says so; the 3.0 run printed no warning, and both treated the pattern the same way:

grep -E '*a' lines.txt | head -2; echo status $?
grep: warning: * at start of expression cat concatenate status 0

Putting It Together

Lines that start with a date and are an error or a warning:

grep -E '^[0-9]{4}-[0-9]{2}-[0-9]{2} (ERROR|WARN)' lines.txt
2026-10-08 ERROR disk full 2026-10-08 WARN low memory

Read it left to right: start of line, four digits, a dash, two digits, a dash, two digits, a space, then either word. (egrep and fgrep are old names for grep -E and grep -F; both ran silently on the two versions here; I understand some newer GNU releases warn that they are obsolete, which I did not see, so write the option instead.)

Hands-On Exercises

Exercise 1

A log has well-formed lines, lines with a badly written date, lowercase levels, indented lines and a WARNING that is not a WARN. With one extended expression print only the lines that start with a full date and time, then ERROR or WARN as a whole word, then a space.

๐Ÿ“„ View solution
Exercise 2

In a paragraph, print each repeated word (ignoring case) with -o, then print every word that has a doubled letter in it, using a back-reference.

๐Ÿ“„ View solution
Exercise 3

Write a small shell function that takes a basic pattern and an extended pattern, runs both on lines.txt with -o, and prints whether they gave the same output. Use it to check four correct basic-to-extended translations and to catch two translations that look right and are not.

๐Ÿ“„ View solution

Chapter 3 Quick Reference

  • Default = basic (BRE), -E = extended (ERE): same ideas, different punctuation. Perl (-P) is Chapter 5
  • Position: ^ start, $ end (^$ = empty line). In basic, $ is an anchor only at the end of the pattern; in extended it is one everywhere
  • . any one character; * zero or more (greedy: .* runs to the last match; x* matches every line)
  • Sets: [abc], [^abc], [a-z]; ] first, specials lose their meaning inside; classes need double brackets: [[:digit:]]
  • Extended: + ? | ( ) {n,m} are operators and the backslash makes them literal; basic is the reverse (GNU: \+ \? \| \( \) \{ \}); a wrong choice is silent
  • \1โ€ฆ\9 repeat what a group matched; \b \< \> \B \w \W \s \S are GNU; \d is NOT a digit here
  • Longest leftmost match wins, whatever the order of alternatives
  • CRLF files: on Linux the carriage return is part of the line and $ fails; Git Bash's grep hides it, so the same pattern can behave differently
  • Invalid pattern: message and status 2; in basic an unmatched ( is just text
Coming next
grep: Searching Text 4 controls what grep prints around the match: context lines with -A, -B and -C, and the output options -n -c -o -b -H -h -m -q -s, with the difference between counting lines and counting matches.