Pipelines and Scripts
grep: Searching Text
Chapter 6 ยท grep in Pipelines and Scripts
At the prompt, grep prints something and you look at it. In a script, nobody is looking: the exit status is the answer, and the next command depends on it. That is where grep's “no match is not an error, but it is not success either” behaviour (Chapter 1) starts to bite. This chapter is mostly the traps, each shown failing and then fixed.
ps example needs a real Linux process
list, so it ran on WSL (3.11) only, and the timing example is compared by seconds, not by exact equality. I did not try macOS/BSD grep, and several scripts use bash features (<<<, <( )), so run them with bash, not plain sh.
Practice Folder
Paste this into an empty folder. Each example below is independent: it assumes this folder, freshly made.
grep as a Question: if grep -q
The usual shape is if grep -q PATTERN FILE. -q keeps the screen clean; the status is the answer.
&& and || read the same status. Be careful with a && b || c as a stand-in for if/else: c runs also if b fails. For anything beyond one short command, use if.
To ask about text you already have in a variable, feed it in with a here-string (<<<, bash) instead of echo "$var" | grep:
Trap 1: set -e and the Status 1
Many scripts start with set -e, which means “stop the script the moment any command fails”. To the shell, grep finding nothing is a failure (status 1), so the script stops at the grep, silently.
Nothing is printed; the script's own exit status is 1:
Three ways out. Add || true (“and if that fails, carry on”), or put the grep in an if condition, where set -e deliberately does not apply:
Pick the first when you do not care about the answer, and the second when you do. The same trap applies to every grep whose result you assign or pipe, which brings us to counting:
Trap 2: Counting Under set -e
grep -c prints 0 and exits with status 1 when nothing matches (Chapter 4). In a command substitution under set -e, that kills the script before it can print the 0:
The fix is || true after the grep. The tempting || echo 0 is wrong, because grep has already printed its own 0, so the variable ends up with two lines:
If you want a count and no fuss, awk counts without the status problem, because awk only reports an error for a real error:
Trap 3: pipefail
Normally a pipeline's status is that of its last command, so a failure in the middle is invisible. set -o pipefail changes that: the pipeline fails if any command in it fails.
Combined with set -e, a grep that finds nothing now stops the script, even though wc at the end succeeded and printed its 0:
Again, || true or an if fixes it, or awk in place of grep.
A stranger one: -q and SIGPIPE
-q stops at the first match. If the command feeding grep is still writing, it gets a broken pipe signal and dies, which pipefail then reports as status 141 (128 plus the signal number 13), even though grep found the answer:
The same pipeline without pipefail, and with pipefail but without -q (grep then reads everything, so nothing is cut off), both say 0:
If the producer is a command, you can also read it through process substitution, which does not take part in the pipeline's status (bash):
A status of 141 is not a match failing; treat it as “grep stopped early”. The cure is any of: no -q, process substitution, or leave pipefail off for that line.
Patterns From Variables
A pattern that arrives in a variable (a user's input, a file name) should be treated as text, not as a regular expression. Use -F, and -- to keep a value that starts with a dash from being read as an option,
and quote the variable. Without -F the dot below matches the x too:
Finding a Process
ps lists every process, and grep filters it, but the grep you are running is also a process whose command line contains the word you are looking for, so it finds itself. (This ran on WSL, with a background
sleep 300 standing in for the process you want.)
The trick is the bracket expression: [s]leep matches the text sleep, but the grep process's own command line contains the literal text [s]leep, which it does not match.
pgrep, where it exists (here, WSL had it and Git Bash did not), does the job without the trick; -x asks for an exact name and -c counts.
Comparing Two Lists
-f FILE takes the patterns from a file. With -F (plain text) and -x (whole line) the patterns become “exactly these lines”, which turns grep into a set tool, and neither file needs sorting.
Our lists: list_a.txt has alice, bob, carol, dave; list_b.txt has bob, dave, erin, bob.
The lines of list_b.txt that are also in list_a.txt. Note that bob appears twice: grep prints the lines of the file being searched as they are, duplicates included. -v turns it into a difference:
That is “in b but not in a”, and the lines of a that are also in b. Swap the files for the other direction. comm does the same job and removes duplicates, but only on sorted input:
-Ff without -x matches everything and the comparison silently becomes meaningless. With -x a blank pattern can only match a blank line:
Find, Then Change: grep and sed
A common job: find every file with a word and change it. The safe version has four habits: preview first, keep backups, handle awkward file names, and never touch the backups themselves. The practice folder has three files that say colour, one of them with a space in its name.
Preview, printing only the lines that would change:
Then apply it, with -i.bak (GNU sed edits in place and keeps the original with a .bak ending), and check:
-lZ and xargs -0 (Chapter 2) kept site/my notes.txt in one piece. Without them it breaks:
sed was asked for site/my and notes.txt, neither of which exists, and it reported that it could not read them. (It still did the files with ordinary names, which is why the message is easy to miss.)
The backups bite back
The .bak files still say colour, so a second run finds them, edits them, and makes backups of the backups:
Keep the backups out of the search with --exclude. Now the second run has nothing to find, which brings up another trap: xargs still runs sed once, with no files, and sed complains:
GNU xargs has -r (--no-run-if-empty) for this: with no input it runs nothing. (BSD xargs, as on macOS, already behaves that way, and I did not try it.) With both fixes, running the edit twice is
harmless:
To undo the change, move each backup back over its file (zero-byte separators again, so spaces are safe):
That last 0 is the count of backups left. Always run on a clean working tree, or a copy: in a project under git, git diff is also a way back, and the .git folder should be excluded
(--exclude-dir=.git) so the edit does not touch it.
grep Next to cut, awk, sort and uniq
grep chooses lines; cut and awk choose fields from them; sort and uniq -c count. Which clients got a 404?
The second picks the GET requests, takes the third field (the path) with awk, and counts how often each was asked for. For fixed layouts like this, awk alone can do the filtering too
(awk '/GET/ {print $3}'), which saves a process, but the grep-then-pick form is easier to read and to change.
Watching a Live Log: --line-buffered
When grep's output goes to a pipe, grep collects it in a buffer and sends it in chunks, so a pipeline that is meant to react immediately (tail -f app.log | grep ERROR | ...) can sit on a line for a long time.
Here the first command sends ERROR1 at once and then waits two seconds before finishing; the last command prints when each line reaches it:
--line-buffered makes grep pass each line on straight away:
The first shows ERROR1 arriving after two seconds, when grep ended and emptied its buffer; the second well before that (the clock counts whole seconds, so what matters is the difference). If a later stage in the pipeline also
buffers (some do) you will still see delays, and the buffering is a performance feature: use --line-buffered when promptness matters more than speed.
Hands-On Exercises
A script starts with set -euo pipefail and counts ERROR and CRITICAL lines in app.log with grep -c. Show that it dies silently before printing the second count, then fix it two ways, with || true and with awk, and say what each fix changes.
Given required.txt and installed.txt, print what is required and missing, what is installed and not required, and what is in both. The required file ends with an accidental blank line: show how that breaks one of the three answers and how to fix it.
In the site folder, rename colour to color safely: preview, apply with backups, confirm, then run it a second time to show nothing more changes, then restore every file from its backup.
Chapter 6 Quick Reference
if grep -q PAT FILE; then ...;if ! grep -q; here-stringgrep -q PAT <<< "$var"(bash)set -e+ a grep with no match (status 1) = the script stops, silently. Fixes:|| true, or use it as anifconditionn=$(grep -c PAT f || true)(not|| echo 0, which gives0twice);awk '/PAT/ {n++} END {print n+0}'never fails on zeropipefail: a grep with no match fails the whole pipeline;producer | grep -qcan give 141 (SIGPIPE) even on a match: drop-q, or use< <(producer)- Variables as patterns:
grep -F -- "$pat"(text, not regex; a value starting with-is safe) ps -eo args | grep '[s]leep'avoids matching grep itself;pgrep -x NAMEwhere available- Lists:
grep -Fxf A B(lines of B that are in A, duplicates kept);grep -vFxf A B(lines of B not in A); a blank pattern line matches everything unless-x - Bulk edit: preview with
sed -n 's/old/new/gp';grep -rlZ OLD --exclude='*.bak' . | xargs -0 -r sed -i.bak 's/OLD/NEW/g'; restore withfind -print0 | while read -d ''andmv --line-bufferedwhen grep's output feeds a pipeline that must react at once