Perl Regex

grep: Searching Text

Chapter 5 · Perl Regex with -P

The basic and extended patterns of Chapter 3 cover most searches. -P switches grep to Perl-compatible regular expressions (PCRE), the same pattern language as Perl, PHP, and (with small differences) Python and JavaScript. It adds things the other two do not have: \d for a digit, matching as little as possible, checking what comes before or after without including it, and throwing away the start of a match. It also has real costs, which this chapter is just as keen on.

Run on two versions of grep
Every example was run on GNU grep 3.0 (Git Bash) and GNU grep 3.11 (WSL). The only differences were the wording of two error messages, shown where they come up. -P is a GNU feature. It worked on both versions here; macOS/BSD grep is often built without it, which I did not try, so check before relying on it in a script that has to run elsewhere.

Practice File

Paste this into an empty folder. It makes text.txt: orders and dates, an HTML line, quoted strings, a tab, and a line with accented letters.

cat > text.txt <<'EOF' order 1042 shipped on 2026-10-08 to bob@example.com order 7 shipped on 2026-09-30 to alice@work.org price: $19.99 (was $25.00) <b>bold</b> and <i>italic</i> text user=bob id=42 role=admin user=alice id=7 role=user foo123bar456 leading spaces here EOF printf 'tabs\there\n' >> text.txt cat >> text.txt <<'EOF' name="Bob" city="Paris" say "hello" and "bye" color colour the the end Price 100 USD, 250 EUR, 75 USD, $30 USD café résumé naïve EOF

Shorthands That Chapter 3 Did Not Have

In Perl mode \d is a digit (Chapter 3 warned that in basic and extended it is not), \s is any white space, \w a word character, and their capitals mean the opposite (\D, \S, \W). \t is a tab and \h is a horizontal space (space or tab). Counts and groups are written the extended way, with no backslashes.

grep -oP '\d{4}-\d{2}-\d{2}' text.txt
2026-10-08 2026-09-30
grep -oP '\w+@\w+\.\w+' text.txt
bob@example.com alice@work.org
grep -P '\t' text.txt | cat -A
tabs^Ihere$
grep -P '^\h' text.txt | cat -A
leading spaces here$

(cat -A shows a tab as ^I and the end of the line as $, so you can see the invisible characters.) None of this needed -P in principle ([0-9] and [[:space:]] do the same job); the shorthands are shorter and read better.

Greedy and Lazy

Chapter 3 showed that .* takes as much as it can. Here is the same trap with HTML. The greedy pattern starts at the first < and runs to the last >:

grep -o '<.*>' text.txt
<b>bold</b> and <i>italic</i>

Add a ? after the repeat and it becomes lazy: it takes as little as it can and still match. That gives each tag separately:

grep -oP '<.*?>' text.txt
<b> </b> <i> </i>

The lazy forms are *?, +?, ?? and {n,m}?. They matter whenever you take a piece out from between two delimiters. The same with quoted strings: greedy swallows from the first quote to the last, lazy takes one string at a time:

grep -o '".*"' text.txt
"Bob" city="Paris" "hello" and "bye"
grep -oP '".*?"' text.txt
"Bob" "Paris" "hello" "bye"

Lazy repeats can also surprise you. \d+ takes a whole number, while \d+? takes as little as is allowed, one digit:

echo 'a1b22c333' | grep -oP '\d+'
1 22 333
echo 'a1b22c333' | grep -oP '\d+?'
1 2 2 3 3 3
You often do not need -P for this
A negated set does the same job in any grep. “A tag” is <, then anything that is not >, then >. The output is identical, and it also works where -P does not exist:
grep -o '<[^>]*>' text.txt
<b> </b> <i> </i>
grep -o '"[^"]*"' text.txt
"Bob" "Paris" "hello" "bye"

Reach for -P and a lazy repeat when the closing delimiter is more than one character, or when a negated set cannot say what you mean.

Looking Around: Lookahead and Lookbehind

A lookaround checks what is next to the match, without making it part of the match. That is exactly what you want with -o: find the number, but only where it is followed by USD, and print just the number.

WriteMeans
X(?=Y)X, when Y follows (lookahead)
X(?!Y)X, when Y does not follow
(?<=Y)XX, when Y comes before (lookbehind)
(?<!Y)XX, when Y does not come before
grep -oP '\d+(?= USD)' text.txt
100 75 30
echo 'Price 100 USD, 250 EUR, 75 USD' | grep -oP '\b\d+\b(?! USD)'
250
grep -oP '(?<=\$)\d+\.\d+' text.txt
19.99 25.00

The first finds 100, 75 and 30 (the numbers followed by USD); the second the one number that is not (250 is in euros); the third the digits after a dollar sign. Combine them. Numbers in USD, but not those that already have a dollar sign in front (the $30 USD at the end):

grep -oP '(?<!\$)\b\d+(?= USD)' text.txt
100 75

Two small traps

A lookbehind must have a fixed length. It has to know how far back to look. \w+ could be any length, so grep refuses:

grep -oP '(?<=\w+@)\w+' text.txt; echo status $?
grep: lookbehind assertion is not fixed length status 2

\K is the way round that. It means “forget everything matched so far”: the text before it must still match, but it is left out of what is printed, with no length restriction:

grep -oP '\w+@\K\w+' text.txt
example work
grep -oP 'id=\K\d+' text.txt
42 7

\K and a lookahead together answer “the user name, but only for the admin”. The name is what we want; the rest of the line only decides:

grep -oP 'user=\K\w+(?=.*role=admin)' text.txt
bob

A lookaround on its own matches nothing visible. (?=foo) matches the position before a foo, with no characters, so the line counts as a match (status 0) but -o has nothing to print:

echo foo | grep -oP '(?=foo)'; echo status $?
status 0

Alternation: First Wins, Not Longest

Chapter 3 said that basic and extended grep take the longest match. Perl mode takes the first alternative that works, so the order you write them in matters:

echo foobar | grep -oP 'foo|foobar'
foo
echo foobar | grep -oE 'foo|foobar'
foobar
echo foobar | grep -oP 'foobar|foo'
foobar

The first is Perl mode, the second the same pattern in extended mode, the third Perl mode with the longer alternative written first. When you move a pattern between -E and -P, check any alternation.

More PCRE Features

(?i) inside the pattern switches case-insensitive matching on from that point (you may know it as an inline flag), which is handy when only part of the pattern should ignore case. A back-reference works as in Chapter 3, and \Q ... \E quotes a stretch of the pattern so that its characters are plain text:

grep -oP '(?i)price' text.txt
price Price
grep -oP '\b(\w+) \1\b' text.txt
the the
grep -oP '\Qprice: $\E\d+' text.txt
price: $19

(?:...) is a group that does not capture (it does not take a number for \1), and PCRE also has named groups and possessive repeats. They are real, but with grep you can only get the whole match out, so they rarely earn their place here.

Limits of -P

One pattern only

-e twice, or -f with several lines, is refused. Join them with | inside the one pattern instead:

grep -P -e 'a' -e 'b' text.txt; echo status $?
grep: the -P option only supports a single pattern status 2

Error messages

An invalid pattern gives a message and status 2, as before. The wording comes from the Perl-regex library and differs by version:

grep -P '(' text.txt; echo status $?
grep: missing ) status 2

That is grep 3.0. In 3.11 the message is grep: missing closing parenthesis. Do not match on the text of a grep error message in a script.

Slow patterns

Perl-style matching works by trying one way and backing up when it fails. A pattern with a repeat inside a repeat, such as (a+)+, can have an enormous number of ways to fail. Here is a line of 30 letters a followed by a b:

head -c 30 /dev/zero | tr '\0' 'a' > aaa.txt; echo b >> aaa.txt
grep -P '^(a+)+$' aaa.txt; echo status $?
grep: exceeded PCRE's backtracking limit status 2

grep gave up with an error (status 2), not a “no match” (status 1). In 3.11 the message is grep: aaa.txt: exceeded PCRE's backtracking limit, with the file name. So an error from -P means “I could not decide”, and a script that treats every non-zero status as “no match” will be wrong. Extended mode answered this particular pattern without trouble:

grep -E '^(a+)+$' aaa.txt; echo status $?
status 1

That answered at once, with a plain “no match”. If you can write the pattern without a repeat inside a repeat (here, simply ^a+$), do. Both of these returned in a fraction of a second here because grep enforces a limit; I did not measure a case without one.

Several Lines at Once: -z

grep matches one line at a time, so a pattern cannot cross a line break. -z changes the separator from a newline to a zero byte, so a text file with no zero bytes becomes one long “line”. Then (?s) lets . match a newline too. The output ends with a zero byte instead of a newline, so tr turns them back. First without (?s), where . stops at the line break and nothing matches:

printf 'start\nmiddle\nend\n' | grep -Pzo 'start.*end' | tr '\0' '\n'
(no output)
printf 'start\nmiddle\nend\n' | grep -Pzo '(?s)start.*?end' | tr '\0' '\n'
start middle end

The lazy .*? ends at the first end; a greedy .* would run to the last one in the file.

Accented Letters

The last line of the practice file is café résumé naïve. In Perl mode \w is ASCII only: it stops at the accented letters and breaks each word into pieces:

tail -1 text.txt | grep -oP '\w+'
caf r sum na ve

Two ways to ask for Unicode letters in PCRE are (*UCP) at the start of the pattern, which makes \w and friends Unicode-aware, and \p{L}, which means “any letter”:

tail -1 text.txt | grep -oP '(*UCP)\w+'
café résumé naïve
tail -1 text.txt | grep -oP '\p{L}+'
café résumé naïve

I confirmed the bytes were identical on both grep versions (WSL's own display of the accents was garbled in my capture, so I compared the output as hex). This depends on a UTF-8 locale; both machines had one. If your text is not English, test your own words before trusting \w.

Choosing Between -E and -P

Use -E when…Use -P when…
the script must run on other machinesyou work on GNU/Linux and need lookaround, \K or lazy repeats
several patterns (-e, -f)one pattern is enough
you want the longest match and no surprisesyou want first-match behaviour, as in Perl, PHP or Python
the pattern has a repeat inside a repeatit does not

Hands-On Exercises

Exercise 1

From the user lines, print the id number of every user, then the name of the user whose role is user (not admin), using \K and a lookahead. Then explain why the same task cannot be done with a lookbehind.

📄 View solution
Exercise 2

A file has the lines name="Bob" city="Paris", say "hello" and "bye" and a line with an escaped quote. Extract each quoted string with a greedy pattern, a lazy one and a negated set, and say which lines each gets right.

📄 View solution
Exercise 3

From a price list, print the numbers in USD that do not have a dollar sign in front, the amounts that are in a currency other than USD, and then show that a pattern with a repeat inside a repeat gives an error rather than “no match”, and rewrite it.

📄 View solution

Chapter 5 Quick Reference

  • -P = Perl-compatible regex: GNU only, one pattern (-e twice is refused: use |), check availability before relying on it
  • \d \D \s \S \w \W \t \h; counts and groups without backslashes; (?i) inline flag; \Q...\E literal text; (?:...) non-capturing
  • Lazy repeats: *? +? ?? {n,m}?; a negated set (<[^>]*>) often does the same job in any grep
  • Lookahead (?=Y) (?!Y), lookbehind (?<=Y) (?<!Y) (fixed length only); \K drops what came before, with no length limit; a bare lookaround matches an empty string, so -o prints nothing
  • Alternation takes the first alternative that works, not the longest (the opposite of -E)
  • -z + (?s) matches across lines; pipe through tr '\0' '\n'
  • \w is ASCII only; (*UCP) or \p{L} for Unicode letters (UTF-8 locale)
  • A repeat inside a repeat can exhaust grep's backtracking limit: that is an error (status 2), not "no match"; error wording differs by version
Coming next
grep: Searching Text 6 puts grep to work in pipelines and scripts: if grep -q, the set -e and pipefail traps, the [n]ginx trick, comparing two lists with -Fxf, and changing files safely after finding them.