Perl Regex
grep: Searching Text
Chapter 5 · Perl Regex with -P
The basic and extended patterns of Chapter 3 cover most searches. -P switches grep to Perl-compatible regular expressions (PCRE), the same pattern language as Perl, PHP, and (with small
differences) Python and JavaScript. It adds things the other two do not have: \d for a digit, matching as little as possible, checking what comes before or after without including it, and throwing away the
start of a match. It also has real costs, which this chapter is just as keen on.
-P is a GNU feature. It worked on both versions here; macOS/BSD grep is often built without it, which I did not try, so check before relying on it in a script that has to run elsewhere.
Practice File
Paste this into an empty folder. It makes text.txt: orders and dates, an HTML line, quoted strings, a tab, and a line with accented letters.
Shorthands That Chapter 3 Did Not Have
In Perl mode \d is a digit (Chapter 3 warned that in basic and extended it is not), \s is any white space, \w a word character, and their capitals mean the opposite
(\D, \S, \W). \t is a tab and \h is a horizontal space (space or tab). Counts and groups are written the extended way, with no backslashes.
(cat -A shows a tab as ^I and the end of the line as $, so you can see the invisible characters.) None of this needed -P in principle ([0-9] and
[[:space:]] do the same job); the shorthands are shorter and read better.
Greedy and Lazy
Chapter 3 showed that .* takes as much as it can. Here is the same trap with HTML. The greedy pattern starts at the first < and runs to the last >:
Add a ? after the repeat and it becomes lazy: it takes as little as it can and still match. That gives each tag separately:
The lazy forms are *?, +?, ?? and {n,m}?. They matter whenever you take a piece out from between two delimiters. The same with quoted strings:
greedy swallows from the first quote to the last, lazy takes one string at a time:
Lazy repeats can also surprise you. \d+ takes a whole number, while \d+? takes as little as is allowed, one digit:
<, then anything that is not >, then >. The output is identical, and it also works where -P does not exist:
Reach for -P and a lazy repeat when the closing delimiter is more than one character, or when a negated set cannot say what you mean.
Looking Around: Lookahead and Lookbehind
A lookaround checks what is next to the match, without making it part of the match. That is exactly what you want with -o: find the number, but only where it is followed by USD,
and print just the number.
| Write | Means |
|---|---|
X(?=Y) | X, when Y follows (lookahead) |
X(?!Y) | X, when Y does not follow |
(?<=Y)X | X, when Y comes before (lookbehind) |
(?<!Y)X | X, when Y does not come before |
The first finds 100, 75 and 30 (the numbers followed by USD); the second the one number that is not (250 is in euros); the third the digits after a dollar sign. Combine them. Numbers in USD, but not those that already
have a dollar sign in front (the $30 USD at the end):
Two small traps
A lookbehind must have a fixed length. It has to know how far back to look. \w+ could be any length, so grep refuses:
\K is the way round that. It means “forget everything matched so far”: the text before it must still match, but it is left out of what is printed, with no length restriction:
\K and a lookahead together answer “the user name, but only for the admin”. The name is what we want; the rest of the line only decides:
A lookaround on its own matches nothing visible. (?=foo) matches the position before a foo, with no characters, so the line counts as a match (status 0) but
-o has nothing to print:
Alternation: First Wins, Not Longest
Chapter 3 said that basic and extended grep take the longest match. Perl mode takes the first alternative that works, so the order you write them in matters:
The first is Perl mode, the second the same pattern in extended mode, the third Perl mode with the longer alternative written first. When you move a pattern between -E and -P, check any alternation.
More PCRE Features
(?i) inside the pattern switches case-insensitive matching on from that point (you may know it as an inline flag), which is handy when only part of the pattern should ignore case. A back-reference works
as in Chapter 3, and \Q ... \E quotes a stretch of the pattern so that its characters are plain text:
(?:...) is a group that does not capture (it does not take a number for \1), and PCRE also has named groups and possessive repeats. They are real, but with grep you can only get the
whole match out, so they rarely earn their place here.
Limits of -P
One pattern only
-e twice, or -f with several lines, is refused. Join them with | inside the one pattern instead:
Error messages
An invalid pattern gives a message and status 2, as before. The wording comes from the Perl-regex library and differs by version:
That is grep 3.0. In 3.11 the message is grep: missing closing parenthesis. Do not match on the text of a grep error message in a script.
Slow patterns
Perl-style matching works by trying one way and backing up when it fails. A pattern with a repeat inside a repeat, such as (a+)+, can have an enormous number of ways to fail. Here is a line of 30
letters a followed by a b:
grep gave up with an error (status 2), not a “no match” (status 1). In 3.11 the message is grep: aaa.txt: exceeded PCRE's backtracking limit, with the file name. So an error from -P means “I could not decide”,
and a script that treats every non-zero status as “no match” will be wrong. Extended mode answered this particular pattern without trouble:
That answered at once, with a plain “no match”. If you can write the pattern without a repeat inside a repeat (here, simply ^a+$), do. Both of these returned in a fraction of a second here because grep enforces a limit;
I did not measure a case without one.
Several Lines at Once: -z
grep matches one line at a time, so a pattern cannot cross a line break. -z changes the separator from a newline to a zero byte, so a text file with no zero bytes becomes one long “line”. Then
(?s) lets . match a newline too. The output ends with a zero byte instead of a newline, so tr turns them back. First without (?s), where . stops at the
line break and nothing matches:
The lazy .*? ends at the first end; a greedy .* would run to the last one in the file.
Accented Letters
The last line of the practice file is café résumé naïve. In Perl mode \w is ASCII only: it stops at the accented letters and breaks each word into pieces:
Two ways to ask for Unicode letters in PCRE are (*UCP) at the start of the pattern, which makes \w and friends Unicode-aware, and \p{L}, which means “any letter”:
I confirmed the bytes were identical on both grep versions (WSL's own display of the accents was garbled in my capture, so I compared the output as hex). This depends on a UTF-8 locale; both machines had one. If your text
is not English, test your own words before trusting \w.
Choosing Between -E and -P
Use -E when… | Use -P when… |
|---|---|
| the script must run on other machines | you work on GNU/Linux and need lookaround, \K or lazy repeats |
several patterns (-e, -f) | one pattern is enough |
| you want the longest match and no surprises | you want first-match behaviour, as in Perl, PHP or Python |
| the pattern has a repeat inside a repeat | it does not |
Hands-On Exercises
From the user lines, print the id number of every user, then the name of the user whose role is user (not admin), using \K and a lookahead. Then explain why the same task cannot be done with a lookbehind.
A file has the lines name="Bob" city="Paris", say "hello" and "bye" and a line with an escaped quote. Extract each quoted string with a greedy pattern, a lazy one and a negated set, and say which lines each gets right.
From a price list, print the numbers in USD that do not have a dollar sign in front, the amounts that are in a currency other than USD, and then show that a pattern with a repeat inside a repeat gives an error rather than “no match”, and rewrite it.
📄 View solutionChapter 5 Quick Reference
-P= Perl-compatible regex: GNU only, one pattern (-etwice is refused: use|), check availability before relying on it\d \D \s \S \w \W \t \h; counts and groups without backslashes;(?i)inline flag;\Q...\Eliteral text;(?:...)non-capturing- Lazy repeats:
*? +? ?? {n,m}?; a negated set (<[^>]*>) often does the same job in any grep - Lookahead
(?=Y) (?!Y), lookbehind(?<=Y) (?<!Y)(fixed length only);\Kdrops what came before, with no length limit; a bare lookaround matches an empty string, so-oprints nothing - Alternation takes the first alternative that works, not the longest (the opposite of
-E) -z+(?s)matches across lines; pipe throughtr '\0' '\n'\wis ASCII only;(*UCP)or\p{L}for Unicode letters (UTF-8 locale)- A repeat inside a repeat can exhaust grep's backtracking limit: that is an error (status 2), not "no match"; error wording differs by version
if grep -q, the set -e and pipefail traps, the [n]ginx trick, comparing two lists with -Fxf, and changing files safely after finding them.