Basic and Extended Regex
grep: Searching Text
Chapter 3 ยท Basic and Extended Regex
So far the patterns were plain words. A pattern can also describe text: “a number”, “a line that starts with a date”, “the same word twice”. That description language is the regular expression, and grep has two spellings of it. This chapter covers grep's own two spellings and where they trip people up. For regular expressions as a subject in their own right (and other tools that use them), see Learning Regular Expressions.
$, and a warning about * at the start of a pattern. Anything marked GNU works in GNU grep (Linux, WSL, Git Bash) and may not in macOS/BSD grep, which I did not try.
One Language, Two Spellings
By default grep reads the pattern as a basic regular expression (BRE). With -E it reads an extended one (ERE). They describe the same things; they differ in
which punctuation characters are operators and which are plain text. Learn the shared core first, then the difference.
Practice File
Paste this into an empty folder. It makes lines.txt, 20 lines chosen to trip up each idea in this chapter.
Where: Anchors
^ matches the start of a line and $ the end. They match a position, not a character.
Only the line that starts with a lowercase the: “The cat sat” has a capital T, and other lines have the in the middle.
Together, ^$ is an empty line, the classic way to count or remove blank lines:
What: One Character and Repeats
. matches any one character. -o prints just what matched, which is the best way to see what a pattern really did:
The last match is a literal c.t (the dot matched a dot). Put * after something to mean “zero or more of it”:
Two traps with *. First, it is greedy: it takes as much as it can. .* (“anything”) between two letters swallows everything up to the last match on the line:
Look at the second line of output: it starts at the first c in concatenate and runs to the last t in the word. Second, “zero or more” includes zero, so a pattern
made only of optional things matches every line, even ones with no x at all:
All 20 lines matched x*. If you mean “one or more”, say so (xx*, or x\+ / x+ below).
Sets: Bracket Expressions
Square brackets match one character from a set. A ^ straight after the [ reverses it, and a dash makes a range:
Inside brackets almost everything stops being special, so to match a literal dot you can write [.] instead of \.. A ] must come first in the set, and a backslash is just a
backslash:
Named classes
Ready-made sets, written with double brackets, which are defined by the locale (the language settings), so they follow that language's idea of a letter or a digit; a range such as [a-z] depends on how the locale orders letters:
| Class | Matches |
|---|---|
[[:digit:]] | 0-9 |
[[:alpha:]] / [[:alnum:]] | letters / letters and digits |
[[:upper:]] / [[:lower:]] | upper / lower case letters |
[[:space:]] | space, tab, newline and similar |
[[:punct:]] | punctuation |
The class name sits inside a bracket expression, so the brackets are doubled. Writing it with one pair is a very common slip, and grep says so:
The Difference: Basic vs Extended
Here is the whole difference. In a basic expression the characters + ? | ( ) { } are ordinary text, and you put a backslash in front of them to make them operators.
In an extended expression (-E) it is the other way round: they are operators, and the backslash makes them ordinary text.
| You want | Basic (default) | Extended (-E) |
|---|---|---|
| one or more | a\+ | a+ |
| optional | a\? | a? |
| either / or | a\|b | a|b |
| group | \(ab\) | (ab) |
| between 2 and 4 | a\{2,4\} | a{2,4} |
| a literal plus sign | a+ | a\+ |
(The backslash forms in the basic column are GNU extensions: they work in GNU grep but not in every grep. Standard basic expressions have no + ? | at all, which is a reason to prefer -E
in scripts that may run elsewhere.) Every one of these has a version that silently does the wrong thing. “Optional u”, first as basic written like extended, which looks for a real question mark:
Status 1, no error, no output: grep was happily looking for the text colou?r. And the mirror image, an extended pattern with a backslash, which is also a literal:
The plus sign is the clearest case, because one line of our file really contains a+b:
The first found the literal text a+b. The second found the line with abab in it, because a+b in extended means “one or more a, then b”. The third asked for a literal plus again.
The same trap with either/or:
-E whenever the pattern needs + ? | ( ) { }, and then write them without backslashes. Reading one set of rules is easier than two. Reserve the basic form for simple patterns with no operators.
Groups, Repeats of Groups and Counts
Brackets repeat one character. A group repeats a whole piece, and {n} or {n,m} says how many times:
Those are the same searches written both ways. The extended ones are shorter and easier to read.
Back-references: Matching Something Again
A group remembers what it matched. \1 means “the same text as group 1 matched”, \2 group 2, and so on. It works in both forms (in -E it is a GNU extension; standard extended expressions do not have it):
That found abab: the group matched ab and \1 demanded ab again. The classic use is finding a repeated word:
\b is a word boundary (below), [a-z]+ is a word, then a space, then \1, the same word again. It found the the, and also
abab abab, whose first abab is a word repeated.
Word Edges and Shorthands (GNU)
| Write | Meaning |
|---|---|
\b | a word boundary: between a word character and a non-word character, or the line edge |
\< and \> | the start and the end of a word |
\B | NOT a word boundary (inside a word) |
\w / \W | a word character (letter, digit, underscore) / anything else |
\s / \S | white space / anything else |
Both found the whole word cat three times (not the one inside concatenate). \B is the reverse, the one inside a word:
\w+ is a neater way to say “a word” than [a-zA-Z0-9_]+:
Note that \d is not “a digit” in grep's basic or extended expressions. It is only a digit in Perl mode (-P, Chapter 5). In basic and extended
it quietly means the letter d; use [0-9] or [[:digit:]].
Matching a Special Character Literally
Put a backslash in front of a character that has a meaning to make it plain, or use -F (Chapter 1) when the whole pattern is plain text. Unescaped, the dot in a file name matches anything:
The dollar sign needs particular care because $ is an anchor in extended expressions wherever it appears, but in a basic expression it is an anchor only at the very
end of the pattern, and plain text anywhere else:
The basic one found the price by accident of that rule; the extended one looked for “the end of a line, then 19”, which can never match. Escape it (\$) and both agree.
Which Match Wins: Longest
When a pattern could match in more than one way, grep takes the longest match at the leftmost position, whatever order you listed the alternatives in:
Perl-style engines (Chapter 5) take the first alternative that works instead, so the same patterns give different answers there. It matters whenever you use -o or edit with the match.
Windows Line Endings
A file written on Windows ends each line with two characters, carriage return and line feed (CRLF). On Linux, grep sees the carriage return as part of the line, so $ (the end of the line) comes after it,
and done$ does not match done:
That is WSL's grep 3.11, behaving as it would on any Linux server. Git Bash's grep 3.0 behaves differently: it is a Windows build and evidently drops the carriage return, so the same first command printed
done and exited with status 0 there, and the second printed nothing. So a pattern that works in Git Bash can fail when the same file reaches a Linux
machine. If a $ pattern mysteriously fails on a file from Windows, look for the carriage returns (cat -A file shows them as ^M), and either match them (as the second command did, using bash's
$'...\r' quoting) or convert the file first.
When grep Complains
A pattern that is not valid gives a message and status 2:
In a basic expression an unmatched ( is only a bracket character, so there is nothing to complain about (status 1, nothing found):
Newer GNU grep also warns about patterns that are legal but usually a mistake. A * at the start of an extended pattern has nothing in front of it to repeat. In 3.11 (shown here) grep says so; the
3.0 run printed no warning, and both treated the pattern the same way:
Putting It Together
Lines that start with a date and are an error or a warning:
Read it left to right: start of line, four digits, a dash, two digits, a dash, two digits, a space, then either word. (egrep and fgrep are old names for grep -E and grep -F;
both ran silently on the two versions here; I understand some newer GNU releases warn that they are obsolete, which I did not see, so write the option instead.)
Hands-On Exercises
A log has well-formed lines, lines with a badly written date, lowercase levels, indented lines and a WARNING that is not a WARN. With one extended expression print only the lines that start with a full date and time, then ERROR or WARN as a whole word, then a space.
In a paragraph, print each repeated word (ignoring case) with -o, then print every word that has a doubled letter in it, using a back-reference.
Write a small shell function that takes a basic pattern and an extended pattern, runs both on lines.txt with -o, and prints whether they gave the same output. Use it to check four correct basic-to-extended translations and to catch two translations that look right and are not.
Chapter 3 Quick Reference
- Default = basic (BRE),
-E= extended (ERE): same ideas, different punctuation. Perl (-P) is Chapter 5 - Position:
^start,$end (^$= empty line). In basic,$is an anchor only at the end of the pattern; in extended it is one everywhere .any one character;*zero or more (greedy:.*runs to the last match;x*matches every line)- Sets:
[abc],[^abc],[a-z];]first, specials lose their meaning inside; classes need double brackets:[[:digit:]] - Extended:
+ ? | ( ) {n,m}are operators and the backslash makes them literal; basic is the reverse (GNU:\+ \? \| \( \) \{ \}); a wrong choice is silent \1โฆ\9repeat what a group matched;\b \< \> \B \w \W \s \Sare GNU;\dis NOT a digit here- Longest leftmost match wins, whatever the order of alternatives
- CRLF files: on Linux the carriage return is part of the line and
$fails; Git Bash's grep hides it, so the same pattern can behave differently - Invalid pattern: message and status 2; in basic an unmatched
(is just text
-A, -B and -C, and the output options -n -c -o -b -H -h -m -q -s, with the difference between counting lines and counting matches.