This chapter uses nothing new. It takes everything so far and shows how it is actually used, on three jobs you will meet: reading a server log, hunting through a codebase, and auditing configuration files. In each, the work
goes the same way: look at the data, ask a narrow question, read the answer, refine the question. The pipelines get long, but each stage is a single idea from an earlier chapter.
Run on two versions of grep
Every example ran on GNU grep 3.0 (Git Bash) and 3.11 (WSL) against the same data. The output was identical except for one thing, which you have met: the binary-file message, which is on stdout in 3.0 and stderr in 3.11 and so appears in a different place in a listing.
All of the data is made up. The “secrets” are the AWS documentation's own example key and invented words; never paste a real secret into a terminal, a chat or a document, which is why these examples print file names and line numbers only.
The Practice Data
Paste this into an empty folder. It makes a request log (service.log, plus one rotated and compressed older log), a small project folder (repo) and a few configuration files (etc).
cat > service.log <<'EOF'
2026-10-08 09:14:02 INFO req=a1f3 GET /index 200 took=35ms
2026-10-08 09:14:05 INFO req=b2c4 GET /about 200 took=22ms
2026-10-08 09:14:09 INFO req=c3d5 POST /login 302 took=180ms
2026-10-08 09:14:41 WARN req=d4e6 GET /search 200 took=1250ms slow query
2026-10-08 09:15:02 INFO req=e5f7 GET /index 200 took=30ms
2026-10-08 09:15:30 ERROR req=f6a8 POST /checkout 500 took=2210ms database timeout
2026-10-08 09:15:31 INFO req=f6a8 retry 1 for checkout
2026-10-08 09:15:34 ERROR req=f6a8 POST /checkout 500 took=3050ms database timeout
2026-10-08 09:16:10 INFO req=a7b9 GET /index 200 took=28ms
2026-10-08 09:16:44 ERROR req=b8c0 GET /missing 404 took=12ms not found
2026-10-08 09:17:03 INFO req=c9d1 GET /about 200 took=25ms
2026-10-08 09:17:50 WARN req=d0e2 GET /search 200 took=1800ms slow query
2026-10-08 09:18:12 ERROR req=e1f3 POST /login 500 took=95ms auth backend down
2026-10-08 09:18:13 ERROR req=e1f3 POST /login 500 took=90ms auth backend down
2026-10-08 09:19:55 INFO req=f2a4 GET /index 200 took=33ms
2026-10-08 09:20:07 INFO req=a3b5 GET /index 200 took=31ms
2026-10-08 09:21:20 ERROR req=b4c6 GET /missing 404 took=10ms not found
2026-10-08 09:22:02 INFO req=c5d7 GET /about 200 took=24ms
2026-10-08 09:25:00 INFO req=d6e8 GET /index 200 took=29ms
2026-10-08 09:31:14 WARN req=e7f9 GET /search 200 took=1510ms slow query
EOF
printf '2026-10-08 08:59:58 ERROR req=z9y8 GET /index 500 took=70ms old error\n2026-10-08 08:59:59 INFO req=z9y9 GET /about 200 took=20ms\n' | gzip > service.log.1.gz
mkdir -p repo/src repo/config repo/tests repo/node_modules/lib repo/dist repo/assets repo/.git
cat > repo/src/app.js <<'EOF'
// TODO(alice): handle timeout
const apiKey = "AKIAIOSFODNN7EXAMPLE";
// FIXME validate input
function start() { return legacyParse(input); }
EOF
cat > repo/src/db.py <<'EOF'
# TODO: use a connection pool
password = "hunter2"
# HACK: retry forever
db_host = os.environ["DB_HOST"]
EOF
cat > repo/src/util.js <<'EOF'
// TODO remove legacyParse once v2 ships
function legacyParse(x) { return x; }
module.exports = { legacyParse };
EOF
cat > repo/config/settings.py <<'EOF'
SECRET_KEY = 'not-really-secret-123'
DEBUG = True
EOF
cat > repo/tests/test_db.py <<'EOF'
# TODO add more tests
password = "test-password"
EOF
cat > repo/node_modules/lib/index.js <<'EOF'
// TODO vendored noise
const password = "vendored";
EOF
cat > repo/id_rsa.example <<'EOF'
-----BEGIN RSA PRIVATE KEY-----
FAKEFAKEFAKEFAKEFAKEFAKEFAKEFAKE
-----END RSA PRIVATE KEY-----
EOF
{ printf '/*TODO minified*/password="min";'; head -c 300 /dev/zero | tr '\0' 'x'; echo; } > repo/dist/app.min.js
printf 'logo\0TODO inside a binary\n' > repo/assets/logo.bin
printf '[core]\n# TODO placeholder in .git\n' > repo/.git/config
mkdir -p etc
cat > etc/sshd_config <<'EOF'
# SSH server configuration
Port 22
PermitRootLogin yes
PasswordAuthentication yes # TODO tighten
#PermitEmptyPasswords no
X11Forwarding no
EOF
cat > etc/app.prod.ini <<'EOF'
; production
debug = false
log_level = info
secret_key = ${SECRET_KEY}
EOF
cat > etc/app.staging.ini <<'EOF'
; staging
debug = true
log_level = debug
EOF
cat > etc/app.dev.ini <<'EOF'
; development
log_level = debug
EOF
The project folder looks like this:
repo/
├── assets/
│ └── logo.bin (a binary file that contains the word TODO)
├── config/
│ └── settings.py
├── dist/
│ └── app.min.js (minified: one very long line)
├── node_modules/
│ └── lib/
│ └── index.js (third-party code)
├── src/
│ ├── app.js
│ ├── db.py
│ └── util.js
├── tests/
│ └── test_db.py
├── .git/
│ └── config
└── id_rsa.example (a fake key file)
Job 1: Reading a Server Log
Each line of service.log has a date, a time, a level, a request id (req=...), the request, a status, how long it took and sometimes a message. The first question is always “what is in here?”:
2026-10-08 09:15:30 ERROR req=f6a8 POST /checkout 500 took=2210ms database timeout
2026-10-08 09:15:34 ERROR req=f6a8 POST /checkout 500 took=3050ms database timeout
2026-10-08 09:16:44 ERROR req=b8c0 GET /missing 404 took=12ms not found
2026-10-08 09:18:12 ERROR req=e1f3 POST /login 500 took=95ms auth backend down
2026-10-08 09:18:13 ERROR req=e1f3 POST /login 500 took=90ms auth backend down
That works when the window lines up with the digits. For an awkward one, such as 09:17:30 to 09:21:10, a pattern gets ugly and awk is better, because fixed-width times compare correctly as text. Here the same answer both ways:
2026-10-08 09:15:30 ERROR req=f6a8 POST /checkout 500 took=2210ms database timeout
2026-10-08 09:15:34 ERROR req=f6a8 POST /checkout 500 took=3050ms database timeout
2026-10-08 09:16:44 ERROR req=b8c0 GET /missing 404 took=12ms not found
2026-10-08 09:18:12 ERROR req=e1f3 POST /login 500 took=95ms auth backend down
2026-10-08 09:18:13 ERROR req=e1f3 POST /login 500 took=90ms auth backend down
Both depend on every line having the same layout. A line with a different date format would slip through either one without a warning.
Follow one request
The request id ties lines together. Request f6a8 failed, was retried and failed again:
grep 'req=f6a8' service.log
2026-10-08 09:15:30 ERROR req=f6a8 POST /checkout 500 took=2210ms database timeout
2026-10-08 09:15:31 INFO req=f6a8 retry 1 for checkout
2026-10-08 09:15:34 ERROR req=f6a8 POST /checkout 500 took=3050ms database timeout
grep -c 'req=f6a8' service.log
3
Note the retry line, which says INFO even though it belongs to a failing request: searching only for ERROR would have missed that context. When ids could be prefixes of longer ids, use -w.
What kinds of thing fail?
Take only the error lines, keep just method, path, status, and count:
2 POST /login 500
2 POST /checkout 500
2 GET /missing 404
Two logins failing with 500, two checkouts with 500 and two 404s. The two kinds of 500 matter more than the 404s, and now they are visible at a glance.
What was slow?
Pull out the number after took=, keep those over a second, and sort (the \K trick from Chapter 5):
2026-10-08 09:14:41 WARN req=d4e6 GET /search 200 took=1250ms slow query
2026-10-08 09:15:30 ERROR req=f6a8 POST /checkout 500 took=2210ms database timeo
2026-10-08 09:15:34 ERROR req=f6a8 POST /checkout 500 took=3050ms database timeo
2026-10-08 09:17:50 WARN req=d0e2 GET /search 200 took=1800ms slow query
2026-10-08 09:31:14 WARN req=e7f9 GET /search 200 took=1510ms slow query
(cut -c1-80 shortens long lines to fit.) Both agree there are five slow requests.
The rotated logs
Old logs are compressed. zgrep is grep for .gz files: it reads the compressed file without unpacking it, and takes the same options. (It was present on both machines.)
The first counts the errors in yesterday's log. The second searches the current and the rotated log together: seven errors in all. -h keeps the file name from being added when there are two files.
Job 2: Hunting Through a Codebase
Searching a project has two problems the log did not: most of the files are not yours (vendored libraries, build output, .git) and some of them are not text. First, the search with no restrictions, with the lines cut short:
grep -rn TODO repo | cut -c1-90
repo/.git/config:2:# TODO placeholder in .git
Binary file repo/assets/logo.bin matches
repo/dist/app.min.js:1:/*TODO minified*/password="min";xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
repo/node_modules/lib/index.js:1:// TODO vendored noise
repo/src/app.js:1:// TODO(alice): handle timeout
repo/src/db.py:1:# TODO: use a connection pool
repo/src/util.js:1:// TODO remove legacyParse once v2 ships
repo/tests/test_db.py:1:# TODO add more tests
Eight hits for TODO. Four are the team's own notes in the source and the tests; the other four are noise: the .git folder, third-party code in node_modules, a binary file and the generated bundle in dist. (Under 3.11 the binary line reads grep: repo/assets/logo.bin: binary file matches and
comes first, because it is on standard error.) Apply the habits from Chapter 2 (-I, --exclude-dir) and look for the markers as whole words:
repo/dist/app.min.js:1:/*TODO minified*/password="min";xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
repo/src/app.js:1:// TODO(alice): handle timeout
repo/src/app.js:3:// FIXME validate input
repo/src/db.py:1:# TODO: use a connection pool
repo/src/db.py:3:# HACK: retry forever
repo/src/util.js:1:// TODO remove legacyParse once v2 ships
repo/tests/test_db.py:1:# TODO add more tests
The first line is the minified file in dist: it is generated, it is one huge line, and it is not worth reading. (This is why cut is at the end of the pipeline: without it that single line would fill the screen.)
Exclude the build folder as well:
repo/src/app.js:1:// TODO(alice): handle timeout
repo/src/app.js:3:// FIXME validate input
repo/src/db.py:1:# TODO: use a connection pool
repo/src/db.py:3:# HACK: retry forever
repo/src/util.js:1:// TODO remove legacyParse once v2 ships
repo/tests/test_db.py:1:# TODO add more tests
Six notes (four TODO, one FIXME and one HACK) in four files. Count them by kind, and by file:
repo/dist/app.min.js:1:TODO minified*/pass
repo/src/app.js:1:TODO(alice): handle
repo/src/db.py:1:TODO: use a connect
repo/src/util.js:1:TODO remove legacyP
repo/tests/test_db.py:1:TODO add more tests
Looking for secrets
A search for secrets finds candidates, not proof. Every search below prints only file and line (cut -d: -f1,2 removes the rest of the line), so the secret itself never lands on your screen or in your shell history's output.
A known key format first, an AWS access key id (AKIA and sixteen capitals and digits):
Four hits, and they are not equally interesting. settings.py is a real candidate. db.py too, though its other line reads the host from the environment, which is the right pattern and does not match. The one in tests is probably a harmless test value,
and the one in dist is the minified bundle again. Someone has to look at each. The pattern also misses a password built another way, so finding nothing proves nothing.
A private key announces itself with a header line. It starts with dashes, so the pattern must be given with -e (Chapter 1), and -l is enough since you want to know which file:
One call in app.js and three mentions in util.js (the comment, the definition and the export), so only one real caller remains. The count per file gets you a to-do list for the clean-up.
What grep -r does not know
grep -r knows nothing about .gitignore, or about which files are tracked by git. That is why we had to name .git, node_modules and dist ourselves. Chapter 8 shows tools that do know.
Job 3: Auditing Configuration Files
Configuration files are mostly comments and blank lines, and the setting you care about may appear in a comment, in a different file or not at all. First, the file as the program sees it: no comments, no blank lines.
grep -vE '^[[:space:]]*(#|$)' etc/sshd_config
Port 22
PermitRootLogin yes
PasswordAuthentication yes # TODO tighten
X11Forwarding no
grep -vE '^[[:space:]]*(#|;|$)' etc/app.prod.ini
debug = false
log_level = info
secret_key = ${SECRET_KEY}
[[:space:]]* allows indentation before the comment mark, and the second pattern adds ; because that file uses semicolons for comments. Notice that the inline comment on the PasswordAuthentication line survived: only lines that start with a comment are removed.
Find a setting everywhere
Anchor at the start of the line (so a comment does not count) and ignore case (the keys are not case-sensitive):
etc/sshd_config:3:PermitRootLogin yes
etc/sshd_config:4:PasswordAuthentication yes # TODO tighten
Both are set to yes, which is the opposite of what you want on a server.
Which files do not set it?
-L lists the files with no match (Chapter 2). The obvious try is misleading:
grep -rL -i debug etc
etc/sshd_config
It lists only sshd_config, but app.dev.ini does not set debug either: it merely contains the word, in log_level = debug. A search for the setting has to say it is the start of a line and followed by an equals sign:
Put the rule in a function and the expectations in a list. For each setting there are three outcomes, and they are different problems: it is right, it is set to the wrong value, or it is not set at all (the default applies):
check() { # check FILE KEY EXPECTED
if grep -qiE "^[[:space:]]*$2[[:space:]=]+$3[[:space:]]*(#.*)?$" "$1"; then
echo "PASS $1: $2 = $3"
elif grep -qiE "^[[:space:]]*$2([[:space:]=]|$)" "$1"; then
echo "FAIL $1: $2 is set, but not to $3"
else
echo "MISSING $1: $2 is not set"
fi
}
check etc/sshd_config PermitRootLogin no
check etc/sshd_config PasswordAuthentication no
check etc/sshd_config X11Forwarding no
check etc/sshd_config PermitEmptyPasswords no
check etc/app.prod.ini debug false
check etc/app.staging.ini debug false
check etc/app.dev.ini debug false
FAIL etc/sshd_config: PermitRootLogin is set, but not to no
FAIL etc/sshd_config: PasswordAuthentication is set, but not to no
PASS etc/sshd_config: X11Forwarding = no
MISSING etc/sshd_config: PermitEmptyPasswords is not set
PASS etc/app.prod.ini: debug = false
FAIL etc/app.staging.ini: debug is set, but not to false
MISSING etc/app.dev.ini: debug is not set
Read the result carefully. PermitEmptyPasswords is MISSING because the only line for it is commented out, and the audit says so rather than passing it. The debug setting passes in production,
fails in staging, and is missing in development. A function like this is the start of a real configuration check, and because it uses grep -q (Chapter 1) it also works in a script: it can count the FAIL lines and exit with a status.
$2 and $3 are put into the pattern as they are, which is fine for simple names and values, and wrong for anything with regular-expression characters in it. (Chapter 6: text and patterns are different things.)
Habits That Carried Through All Three
Build the pipeline one stage at a time. Run the grep, read its output, then add the next stage. A long pipeline that does not work is much harder to debug than a short one that does.
Narrow the data before narrowing the pattern.--exclude-dir, -I, --include and anchoring at the start of the line removed more noise than any clever pattern did.
Count, then look.uniq -c and -c tell you whether a problem is one line or two hundred before you read any of them.
Treat the result as candidates. Searches find what they are told to find. The test-file password and the minified bundle were both “correct” hits and neither was a problem.
Keep secrets off the screen. Print where, not what.
Save what worked. The audit function is more useful as a file in your scripts folder than as something you will retype next month.
Hands-On Exercises
Exercise 1
From service.log, list the three slowest requests with their request id, method, path and time, slowest first. Then count the errors in the current and the rotated log together and show how many of the requests in the rotated log failed.
Produce a two-part report for repo: how many TODO, FIXME and HACK notes each file has (ignoring .git, node_modules and dist), and which files hold probable secrets, as file and line number only. Then explain why two of the secret candidates are probably not a problem.
Extend the audit: add two more checks of your choice on etc/sshd_config, make the script print a one-line summary of the number of FAIL and MISSING results, and exit with status 1 if there are any.
Logs: grep -oE 'INFO|WARN|ERROR' | sort | uniq -c | sort -rn; time windows with a pattern (09:1[5-8]) or awk '$2 >= "09:15:00" && $2 <= "09:18:59"'; follow an id with grep 'req=ID'; slow requests grep -oP 'took=\K[0-9]+' | awk '$1 > 1000'; rotated logs with zgrep
Codebase: grep -rnIE --exclude-dir=.git --exclude-dir=node_modules '\b(TODO|FIXME|HACK)\b' .; exclude generated folders too; cut -c1-N tames long lines; -w and --include to find a function's uses
Secrets: known formats (AKIA[0-9A-Z]{16}), assignments (password *= *"..."), key headers (with -e); print file:line only (cut -d: -f1,2, or -l); results are candidates, not proof
Config: grep -vE '^[[:space:]]*(#|$)' shows the effective file (inline comments stay); anchor settings at line start; -L needs a real setting pattern, not just the word; audit as PASS / FAIL / MISSING
Habits: one stage at a time; narrow the data first; count, then look; keep secrets off the screen; save what worked
Coming next
grep: Searching Text 8 is the last chapter: tools that know about git and ignore files (git grep, ripgrep, ag), what each does differently, grep on Windows, and a recap of the whole course.