Context and Output
grep: Searching Text
Chapter 4 ยท Context and Output Control
A match on its own is often not enough. You want the lines around it, which line it was, which file it came from, how many there were, or only the matching part. This chapter is the set of options that shape what grep prints, and the ones that make it quiet.
Two Practice Files
Paste this into an empty folder. app.log has 16 lines: five ERROR matches on four lines (line 14 has the word twice), two WARN lines and the rest INFO.
access.log is a small web server log.
Lines Around the Match: -A, -B and -C
An error is rarely understood on its own line. -A N prints N lines After each match, -B N N lines Before, and -C N N lines of
Context on both sides:
-NUM is a shorthand for -C NUM: -2 means two lines either side.
Groups, separators and overlap
When two matches are close, their context overlaps and grep prints the lines once, as one group. When there is a gap, it prints a line containing -- between the groups so you can see
where one ended and the next began. Here, the three ERROR lines at 5, 7 and 8 became one group, and line 14 is a second:
With -n (line numbers) the output shows which lines are matches and which are context. A colon follows the number of a match, a dash follows the number of a context line:
The same rule applies when the file name is shown, so you can tell matches from context in a long listing. It also means a file name that itself contains a dash can be ambiguous to read:
You can change the separator, or remove it:
-o and context do not mix well: -o prints only the matching text, so there is no context to show, but grep still prints the -- separators between groups. And -c ignores context
completely (the count is of matching lines only):
Which Line, Which File, How Far In
-n adds the line number:
-b adds the byte offset, how many bytes into the file the line starts. It is rarely useful for people, but it is how you would jump to a spot with other tools. Combine with -n and you get
line number, then offset:
-H always shows the file name, and -h never does. By default grep shows it only when there is more than one file to search (Chapter 1). -H is what you want when you search one file but
another tool will read the output, and -h when you search many files but want only the lines:
When the input is a pipe there is no file name at all, so -H shows (standard input). --label=NAME lets you say what to call it:
Only the Match: -o
-o prints just the part of the line that matched, one match per line of output, even when one input line has several. Line 14 contains ERROR twice, so it appears twice:
The byte offset with -o is the position of the match itself, not of the line, which is how -ob tells two matches on the same line apart:
Counting: Lines or Matches
-c prints a count of lines that match, not of matches:
There are five ERROR matches (see the -o output above) but only four lines, because line 14 counts once. To count matches, print one per line with -o and let
wc -l count them:
With several files, -c gives one count per file, with the name in front, and prints 0 for a file with no match. -h drops the names:
-v counts the lines that do not match, so -vc is “how many lines are not INFO”. And counting an empty pattern, which matches every line, is a way to count lines:
Stopping Early: -m
-m N stops after N matching lines. The manual says grep also stops reading the file at that point, so it should be faster on a huge file when you only need the first few (I did not time it). Context after the last match is still printed:
Quiet and Silent: -q and -s
-q prints nothing and only sets the exit status (Chapter 1). The manual says it exits immediately on the first match, which should make it the quickest yes-or-no test (not timed here). What I did check is that a match wins even if another file was missing:
-s (silent) hides the error messages about missing or unreadable files. Compare the two runs below, where one file does not exist:
The messages are gone, but the exit status is not changed: a missing file is still an error (status 2). And -s hides only file errors. A bad pattern still complains:
A script that wants a count and wants to know whether the file could be read can use both. With a missing file, -c prints nothing at all (not 0), so the status is the only clue:
(2>/dev/null also hides the messages, but it hides everything grep says on standard error, including a bad pattern, so -s is the more precise choice.)
Colour
--color=auto highlights the match when the output goes to a terminal, and does nothing when it goes to a pipe or a file, so it is safe to leave on. --color=always forces the colour codes
in, even when the output is not a terminal. Piped through cat -v, which shows the invisible characters, here is the difference:
The ^[[01;31m parts are the escape codes. In a terminal they make the word red; in a file or another command's input they are junk that can break the next step. Use --color=auto
(many people put alias grep='grep --color=auto' in their shell startup file) and avoid always except when you want colour through a pager such as less -R.
Counting Things: -o, sort and uniq -c
The most useful combination in the chapter. -o pulls out the thing you care about, sort groups equal items together, uniq -c counts each group, and sort -rn puts the
biggest first. How many of each level are in the log?
(Ten INFO, five ERROR, two WARN. The five ERROR is matches, as the -o counted them, not the four lines.) Which clients made the most requests?
And what status codes did the server return? Here there is no need to sort by size:
uniq only joins neighbouring equal lines, which is why sort comes first. Leave it out and the counts will be wrong with no error to warn you.
Hands-On Exercises
Show every ERROR in app.log with one line before and two lines after, line numbers on, and a separator of -----. Then say which printed lines are matches and which are context, and why there are two groups, not four.
From app.log report, labelled: the number of lines with ERROR, the number of ERROR matches, the number of lines without ERROR, and how many of each level there are, biggest first.
Write errcount.sh FILE that prints how many lines contain ERROR, saying so plainly for a file with none and exiting with 2 (and a message of its own) for a file that cannot be read. Try it on app.log, access.log and a missing file. Then list the two busiest client addresses in access.log.
Chapter 4 Quick Reference
-A Nafter,-B Nbefore,-C Nor-Nboth sides; overlapping context is printed once; non-adjacent groups are separated by--(--group-separator=TEXT,--no-group-separator)- With
-nor file names::marks a matching line,-marks a context line -owith context prints the--separators but no context;-cignores context-nline number;-bbyte offset of the line (of the match itself with-o);-H/-halways / never show the file name;--label=NAMEnames standard input-o: one output line per match;-c: count of lines (use-o | wc -lto count matches);-cprints0for a file with no match and nothing for an unreadable one-m N: stop after N matching lines (context after the last one is still printed; grep stops reading)-qstatus only, exits at the first match, a match beats a missing file;-shides file errors only (status unchanged, a bad pattern still reports)--color=autois safe;alwaysputs escape codes into pipes and files- Counting things:
grep -o PAT | sort | uniq -c | sort -rn;uniqneeds sorted input
-P: \d, lazy matching, lookahead and lookbehind, \K, and where it is not available.