Context and Output

grep: Searching Text

Chapter 4 ยท Context and Output Control

A match on its own is often not enough. You want the lines around it, which line it was, which file it came from, how many there were, or only the matching part. This chapter is the set of options that shape what grep prints, and the ones that make it quiet.

Run on two versions of grep
Every example was run on GNU grep 3.0 (Git Bash) and GNU grep 3.11 (WSL), on the same two practice files. The output was identical on both for every example in this chapter.

Two Practice Files

Paste this into an empty folder. app.log has 16 lines: five ERROR matches on four lines (line 14 has the word twice), two WARN lines and the rest INFO. access.log is a small web server log.

cat > app.log <<'EOF' 09:00:01 INFO service starting 09:00:02 INFO loading config 09:00:03 WARN config key 'timeout' missing, using default 09:00:04 INFO connecting to database 09:00:09 ERROR database connection refused 09:00:10 INFO retrying in 5s 09:00:15 ERROR database connection refused 09:00:16 ERROR giving up after 2 attempts 09:00:17 INFO switching to read-only mode 09:00:20 INFO listening on port 8080 09:05:00 INFO health check ok 09:10:00 WARN slow response: 2300 ms 09:10:30 INFO health check ok 09:15:00 ERROR disk usage 95%, ERROR threshold 90% 09:15:01 INFO alert sent 09:20:00 INFO health check ok EOF cat > access.log <<'EOF' 10.0.0.5 GET /index.html 200 10.0.0.7 GET /about.html 200 10.0.0.5 POST /login 302 10.0.0.9 GET /missing 404 10.0.0.5 GET /index.html 200 10.0.0.7 GET /index.html 200 10.0.0.5 GET /dashboard 200 10.0.0.9 GET /missing 404 10.0.0.7 GET /logout 302 EOF

Lines Around the Match: -A, -B and -C

An error is rarely understood on its own line. -A N prints N lines After each match, -B N N lines Before, and -C N N lines of Context on both sides:

grep -A1 'giving up' app.log
09:00:16 ERROR giving up after 2 attempts 09:00:17 INFO switching to read-only mode
grep -B1 'giving up' app.log
09:00:15 ERROR database connection refused 09:00:16 ERROR giving up after 2 attempts
grep -B1 -A2 'giving up' app.log
09:00:15 ERROR database connection refused 09:00:16 ERROR giving up after 2 attempts 09:00:17 INFO switching to read-only mode 09:00:20 INFO listening on port 8080

-NUM is a shorthand for -C NUM: -2 means two lines either side.

grep -2 'giving up' app.log
09:00:10 INFO retrying in 5s 09:00:15 ERROR database connection refused 09:00:16 ERROR giving up after 2 attempts 09:00:17 INFO switching to read-only mode 09:00:20 INFO listening on port 8080

Groups, separators and overlap

When two matches are close, their context overlaps and grep prints the lines once, as one group. When there is a gap, it prints a line containing -- between the groups so you can see where one ended and the next began. Here, the three ERROR lines at 5, 7 and 8 became one group, and line 14 is a second:

grep -C1 ERROR app.log
09:00:04 INFO connecting to database 09:00:09 ERROR database connection refused 09:00:10 INFO retrying in 5s 09:00:15 ERROR database connection refused 09:00:16 ERROR giving up after 2 attempts 09:00:17 INFO switching to read-only mode -- 09:10:30 INFO health check ok 09:15:00 ERROR disk usage 95%, ERROR threshold 90% 09:15:01 INFO alert sent

With -n (line numbers) the output shows which lines are matches and which are context. A colon follows the number of a match, a dash follows the number of a context line:

grep -n -C1 ERROR app.log
4-09:00:04 INFO connecting to database 5:09:00:09 ERROR database connection refused 6-09:00:10 INFO retrying in 5s 7:09:00:15 ERROR database connection refused 8:09:00:16 ERROR giving up after 2 attempts 9-09:00:17 INFO switching to read-only mode -- 13-09:10:30 INFO health check ok 14:09:15:00 ERROR disk usage 95%, ERROR threshold 90% 15-09:15:01 INFO alert sent

The same rule applies when the file name is shown, so you can tell matches from context in a long listing. It also means a file name that itself contains a dash can be ambiguous to read:

grep -H -n -C1 'giving up' app.log
app.log-7-09:00:15 ERROR database connection refused app.log:8:09:00:16 ERROR giving up after 2 attempts app.log-9-09:00:17 INFO switching to read-only mode

You can change the separator, or remove it:

grep -C1 WARN app.log
09:00:02 INFO loading config 09:00:03 WARN config key 'timeout' missing, using default 09:00:04 INFO connecting to database -- 09:05:00 INFO health check ok 09:10:00 WARN slow response: 2300 ms 09:10:30 INFO health check ok
grep -C1 --no-group-separator WARN app.log
09:00:02 INFO loading config 09:00:03 WARN config key 'timeout' missing, using default 09:00:04 INFO connecting to database 09:05:00 INFO health check ok 09:10:00 WARN slow response: 2300 ms 09:10:30 INFO health check ok
grep -C1 --group-separator='=====' WARN app.log
09:00:02 INFO loading config 09:00:03 WARN config key 'timeout' missing, using default 09:00:04 INFO connecting to database ===== 09:05:00 INFO health check ok 09:10:00 WARN slow response: 2300 ms 09:10:30 INFO health check ok
Context options and other options
-o and context do not mix well: -o prints only the matching text, so there is no context to show, but grep still prints the -- separators between groups. And -c ignores context completely (the count is of matching lines only):
grep -o -C1 WARN app.log
WARN -- WARN
grep -c -C1 ERROR app.log
4

Which Line, Which File, How Far In

-n adds the line number:

grep -n ERROR app.log
5:09:00:09 ERROR database connection refused 7:09:00:15 ERROR database connection refused 8:09:00:16 ERROR giving up after 2 attempts 14:09:15:00 ERROR disk usage 95%, ERROR threshold 90%

-b adds the byte offset, how many bytes into the file the line starts. It is rarely useful for people, but it is how you would jump to a spot with other tools. Combine with -n and you get line number, then offset:

grep -b 'giving up' app.log
275:09:00:16 ERROR giving up after 2 attempts
grep -nb 'giving up' app.log
8:275:09:00:16 ERROR giving up after 2 attempts

-H always shows the file name, and -h never does. By default grep shows it only when there is more than one file to search (Chapter 1). -H is what you want when you search one file but another tool will read the output, and -h when you search many files but want only the lines:

grep -H 'giving up' app.log
app.log:09:00:16 ERROR giving up after 2 attempts
grep -h GET access.log | head -2
10.0.0.5 GET /index.html 200 10.0.0.7 GET /about.html 200

When the input is a pipe there is no file name at all, so -H shows (standard input). --label=NAME lets you say what to call it:

cat app.log | grep -H 'giving up'
(standard input):09:00:16 ERROR giving up after 2 attempts
cat app.log | grep -H --label=app.log 'giving up'
app.log:09:00:16 ERROR giving up after 2 attempts

Only the Match: -o

-o prints just the part of the line that matched, one match per line of output, even when one input line has several. Line 14 contains ERROR twice, so it appears twice:

grep -on ERROR app.log
5:ERROR 7:ERROR 8:ERROR 14:ERROR 14:ERROR

The byte offset with -o is the position of the match itself, not of the line, which is how -ob tells two matches on the same line apart:

grep -ob ERROR app.log
168:ERROR 241:ERROR 284:ERROR 507:ERROR 529:ERROR

Counting: Lines or Matches

-c prints a count of lines that match, not of matches:

grep -c ERROR app.log
4

There are five ERROR matches (see the -o output above) but only four lines, because line 14 counts once. To count matches, print one per line with -o and let wc -l count them:

grep -o ERROR app.log | wc -l
5

With several files, -c gives one count per file, with the name in front, and prints 0 for a file with no match. -h drops the names:

grep -c ERROR app.log access.log
app.log:4 access.log:0
grep -ch ERROR app.log access.log
4 0

-v counts the lines that do not match, so -vc is “how many lines are not INFO”. And counting an empty pattern, which matches every line, is a way to count lines:

grep -vc INFO app.log
6
grep -c '' app.log
16

Stopping Early: -m

-m N stops after N matching lines. The manual says grep also stops reading the file at that point, so it should be faster on a huge file when you only need the first few (I did not time it). Context after the last match is still printed:

grep -m2 ERROR app.log
09:00:09 ERROR database connection refused 09:00:15 ERROR database connection refused
grep -m1 -A1 ERROR app.log
09:00:09 ERROR database connection refused 09:00:10 INFO retrying in 5s
grep -m2 -c ERROR app.log
2

Quiet and Silent: -q and -s

-q prints nothing and only sets the exit status (Chapter 1). The manual says it exits immediately on the first match, which should make it the quickest yes-or-no test (not timed here). What I did check is that a match wins even if another file was missing:

grep -q ERROR app.log nosuch; echo status $?
status 0

-s (silent) hides the error messages about missing or unreadable files. Compare the two runs below, where one file does not exist:

grep ERROR nosuch app.log
grep: nosuch: No such file or directory app.log:09:00:09 ERROR database connection refused app.log:09:00:15 ERROR database connection refused app.log:09:00:16 ERROR giving up after 2 attempts app.log:09:15:00 ERROR disk usage 95%, ERROR threshold 90%
grep -s ERROR nosuch app.log
app.log:09:00:09 ERROR database connection refused app.log:09:00:15 ERROR database connection refused app.log:09:00:16 ERROR giving up after 2 attempts app.log:09:15:00 ERROR disk usage 95%, ERROR threshold 90%

The messages are gone, but the exit status is not changed: a missing file is still an error (status 2). And -s hides only file errors. A bad pattern still complains:

grep -s ERROR nosuch; echo status $?
status 2
grep -sE '(' app.log; echo status $?
grep: Unmatched ( or \( status 2

A script that wants a count and wants to know whether the file could be read can use both. With a missing file, -c prints nothing at all (not 0), so the status is the only clue:

n=$(grep -sc ERROR nosuch); echo "[$n] status $?"
[] status 2

(2>/dev/null also hides the messages, but it hides everything grep says on standard error, including a bad pattern, so -s is the more precise choice.)

Colour

--color=auto highlights the match when the output goes to a terminal, and does nothing when it goes to a pipe or a file, so it is safe to leave on. --color=always forces the colour codes in, even when the output is not a terminal. Piped through cat -v, which shows the invisible characters, here is the difference:

grep --color=auto ERROR app.log | head -1 | cat -v
09:00:09 ERROR database connection refused
grep --color=always -n ERROR app.log | head -1 | cat -v
^[[32m^[[K5^[[m^[[K^[[36m^[[K:^[[m^[[K09:00:09 ^[[01;31m^[[KERROR^[[m^[[K database connection refused

The ^[[01;31m parts are the escape codes. In a terminal they make the word red; in a file or another command's input they are junk that can break the next step. Use --color=auto (many people put alias grep='grep --color=auto' in their shell startup file) and avoid always except when you want colour through a pager such as less -R.

Counting Things: -o, sort and uniq -c

The most useful combination in the chapter. -o pulls out the thing you care about, sort groups equal items together, uniq -c counts each group, and sort -rn puts the biggest first. How many of each level are in the log?

grep -oE 'INFO|WARN|ERROR' app.log | sort | uniq -c | sort -rn
10 INFO 5 ERROR 2 WARN

(Ten INFO, five ERROR, two WARN. The five ERROR is matches, as the -o counted them, not the four lines.) Which clients made the most requests?

grep -oE '^[0-9.]+' access.log | sort | uniq -c | sort -rn
4 10.0.0.5 3 10.0.0.7 2 10.0.0.9

And what status codes did the server return? Here there is no need to sort by size:

grep -oE ' [0-9]{3}$' access.log | sort | uniq -c
5 200 2 302 2 404

uniq only joins neighbouring equal lines, which is why sort comes first. Leave it out and the counts will be wrong with no error to warn you.

Hands-On Exercises

Exercise 1

Show every ERROR in app.log with one line before and two lines after, line numbers on, and a separator of -----. Then say which printed lines are matches and which are context, and why there are two groups, not four.

๐Ÿ“„ View solution
Exercise 2

From app.log report, labelled: the number of lines with ERROR, the number of ERROR matches, the number of lines without ERROR, and how many of each level there are, biggest first.

๐Ÿ“„ View solution
Exercise 3

Write errcount.sh FILE that prints how many lines contain ERROR, saying so plainly for a file with none and exiting with 2 (and a message of its own) for a file that cannot be read. Try it on app.log, access.log and a missing file. Then list the two busiest client addresses in access.log.

๐Ÿ“„ View solution

Chapter 4 Quick Reference

  • -A N after, -B N before, -C N or -N both sides; overlapping context is printed once; non-adjacent groups are separated by -- (--group-separator=TEXT, --no-group-separator)
  • With -n or file names: : marks a matching line, - marks a context line
  • -o with context prints the -- separators but no context; -c ignores context
  • -n line number; -b byte offset of the line (of the match itself with -o); -H / -h always / never show the file name; --label=NAME names standard input
  • -o: one output line per match; -c: count of lines (use -o | wc -l to count matches); -c prints 0 for a file with no match and nothing for an unreadable one
  • -m N: stop after N matching lines (context after the last one is still printed; grep stops reading)
  • -q status only, exits at the first match, a match beats a missing file; -s hides file errors only (status unchanged, a bad pattern still reports)
  • --color=auto is safe; always puts escape codes into pipes and files
  • Counting things: grep -o PAT | sort | uniq -c | sort -rn; uniq needs sorted input
Coming next
grep: Searching Text 5 introduces Perl-style patterns with -P: \d, lazy matching, lookahead and lookbehind, \K, and where it is not available.