Faster Alternatives

grep: Searching Text

Chapter 8 · Faster Alternatives

grep is installed almost everywhere and does one thing well. For searching a project, though, other tools have the sensible defaults built in: they know about git, skip the files you would have excluded by hand, and are noticeably faster. This chapter looks at git grep and ripgrep (rg), and then at what to use on Windows when grep is not around.

What was and was not tried
git grep ran on both git 2.49.0 (Git Bash on Windows) and git 2.43.0 (WSL), with the same lines. Two listings are in a different order (the two sort programs put capital letters in different places), and the “not a git repository” message is longer in 2.43; nothing else differed. ripgrep 15.0.0 is not installed on this machine: I used the copy that ships inside VS Code, from a scratch folder, and installed nothing, so ripgrep ran on Windows only (its paths come out with backslashes). The Silver Searcher (ag) is not on this machine and I did not run it; the little I say about it is from its documentation. PowerShell is Windows PowerShell 5.1. I did not try macOS.

A Practice Repository

Paste this into an empty folder. It creates a small git repository, proj, with two commits, and then adds files that git does not track. The git commands are given your user name locally, so nothing outside the folder is changed.

mkdir -p proj/src proj/build proj/data cd proj git init -q git config user.name Demo git config user.email demo@example.com printf 'build/\n*.log\n.env\n' > .gitignore printf '// TODO(alice): handle timeout\n// FIXME validate input\nfunction start() { return legacyParse(input); }\n' > src/app.js printf '# TODO: use a connection pool\n# TODO old note about caching\ndef connect():\n # HACK: retry forever\n return None\n' > src/db.py printf '# Demo\nTODO write the docs\n' > README.md printf 'logo\0TODO inside a binary\n' > data/logo.bin git add -A git commit -q -m "first" printf '# TODO: use a connection pool\ndef connect():\n # HACK: retry forever\n return None\n' > src/db.py git commit -q -am "drop the old caching note" printf '// TODO generated\n' > build/out.js printf 'ERROR TODO in a log\n' > debug.log printf 'TODO=unset\n' > .env printf 'TODO untracked note\n' > scratch.txt printf 'TODO in a hidden file\n' > .hidden-notes

What you now have, as git sees it:

FileState
src/app.js, src/db.py, README.md, data/logo.bintracked (committed)
build/out.js, debug.log, .envignored, by .gitignore
scratch.txt, .hidden-notesnot tracked and not ignored
git status --short
?? .hidden-notes ?? scratch.txt

(git status lists only the two untracked files: ignored files are not shown.) Every file above contains the word TODO. The earlier commit also had a line, TODO old note about caching, that the second commit removed.

What Plain grep Finds

With the options from Chapter 2 (-r, -I), grep finds every file that contains TODO:

grep -rlI TODO . | sort
./.env ./.git/hooks/sendemail-validate.sample ./.hidden-notes ./build/out.js ./debug.log ./README.md ./scratch.txt ./src/app.js ./src/db.py

Nine files, and only three of them are files git tracks. The rest are the files .gitignore says you do not care about, the hidden file, the untracked scratch note, and some sample scripts inside the .git folder (git puts them there when a repository is created). Chapter 7 dealt with this by naming each folder to skip. These tools do it from what git already knows.

git grep

git grep is part of git, so it is there wherever git is. It searches the files git tracks:

git grep -l TODO
README.md data/logo.bin src/app.js src/db.py

Four files: the three text files and the binary. No .git folder, no ignored files, no untracked ones. It takes many of grep's options (here I used -n -l -e -I -q and -P, which worked on both machines; the manual lists the rest):

git grep -n TODO
README.md:2:TODO write the docs Binary file data/logo.bin matches src/app.js:1:// TODO(alice): handle timeout src/db.py:1:# TODO: use a connection pool
git grep -n -P 'TODO\(\w+\)'
src/app.js:1:// TODO(alice): handle timeout

The binary file is reported with git's own message, the same on both versions. -I skips it, as in grep, and the exit status is 1 when nothing is found (as in grep):

git grep -n TODO -- data; git grep -n -I TODO -- data; echo "status $?"
Binary file data/logo.bin matches status 1

Untracked files, and files that are ignored

A file you have just made but not yet added is not searched. --untracked adds the files that are neither tracked nor ignored, which is usually what you want while you work:

git grep -n --untracked TODO
.hidden-notes:1:TODO in a hidden file README.md:2:TODO write the docs Binary file data/logo.bin matches scratch.txt:1:TODO untracked note src/app.js:1:// TODO(alice): handle timeout src/db.py:1:# TODO: use a connection pool

--no-index goes the other way: it searches the working folder the way grep -r would (ignored files included, .git still left alone):

git grep -l --no-index TODO | sort
.env .hidden-notes build/out.js data/logo.bin debug.log README.md scratch.txt src/app.js src/db.py

Old versions, and the working tree

Name a commit and git grep searches the files as they were then. The note that was removed in the second commit is still in the first:

git grep -n 'old note' HEAD~1
HEAD~1:src/db.py:2:# TODO old note about caching
git grep -n 'old note'; echo "status $?"
status 1

By default it reads the files in your working folder, including changes you have not committed or staged. --cached searches what has been staged. I added a line to README.md without staging it:

echo 'TODO uncommitted' >> README.md echo 'working tree:'; git grep -n uncommitted echo 'index (--cached):'; git grep -n --cached uncommitted; echo "status $?"
working tree: README.md:3:TODO uncommitted index (--cached): status 1

Combining conditions, and choosing paths

--and, --or and --not combine patterns on the same line, which in plain grep takes two greps in a pipe:

git grep -n -e TODO --and -e alice
src/app.js:1:// TODO(alice): handle timeout
git grep -n -e TODO --and --not -e alice -- src
src/db.py:1:# TODO: use a connection pool

After -- come pathspecs, git's own file patterns. '*.js' picks files by name and ':!...' removes some:

git grep -n TODO -- '*.js'
src/app.js:1:// TODO(alice): handle timeout
git grep -n TODO -- ':!*.md' ':!*.bin'
src/app.js:1:// TODO(alice): handle timeout src/db.py:1:# TODO: use a connection pool

Run from a subfolder, it searches that folder and prints paths relative to it:

cd src && git grep -n TODO
app.js:1:// TODO(alice): handle timeout db.py:1:# TODO: use a connection pool

Outside a repository

git grep needs a repository, and says so with status 128 (not grep's 2). --no-index makes it work anywhere:

cd .. git grep -n TODO; echo "status $?" git grep --no-index -n TODO proj/src
fatal: not a git repository (or any of the parent directories): .git status 128 proj/src/app.js:1:// TODO(alice): handle timeout proj/src/db.py:1:# TODO: use a connection pool

ripgrep (rg)

ripgrep searches the current folder recursively by default, and by default it skips what you would usually skip: files matched by .gitignore (its documentation names other ignore files too), hidden files and folders (so .git too), and binary files. Here is the same search with no options at all:

rg TODO | sort
README.md:TODO write the docs scratch.txt:TODO untracked note src\app.js:// TODO(alice): handle timeout src\db.py:# TODO: use a connection pool

The same files as git grep --untracked, except two: rg skipped .hidden-notes because it is hidden, and the binary file because it is binary. scratch.txt was found although git does not track it, because rg works from ignore files, not from git's list of tracked files. Three things about this output:

  • It looks like grep's, file:text. That is the form when the output goes to a pipe or a file. On a terminal rg instead groups results under a file heading, with line numbers and colour. I could only capture piped output here, so I have not shown the terminal form.
  • Line numbers are off in that form: add -n.
  • The files come out in any order, because rg searches several at once. I sorted the output above with | sort (Chapter 4's habit again).
rg -n TODO | sort
README.md:2:TODO write the docs scratch.txt:1:TODO untracked note src\app.js:1:// TODO(alice): handle timeout src\db.py:1:# TODO: use a connection pool

On Windows, rg printed paths with a backslash, as in src\app.js (I expect slashes on Linux and macOS but could not check).

The -u switches: how much to skip

--files lists the files rg would search, without searching: the best way to see what its defaults leave out.

rg --files | sort
data\logo.bin README.md scratch.txt src\app.js src\db.py

Each -u (“unrestricted”) switches off one more kind of skipping: -u stops honouring ignore files, -uu also searches hidden files and folders, -uuu also searches binary files. The number of files grows at every step (mostly .git internals at -uu):

for level in "" -u -uu -uuu; do printf "rg %-5s --files: " "$level"; rg $level --files | wc -l; done
rg --files: 5 rg -u --files: 7 rg -uu --files: 46 rg -uuu --files: 46

With -uuu a binary match is reported, with a note of where the zero byte was:

rg -uuu -n TODO | sort | cut -c1-90
.env:1:TODO=unset .git\hooks\sendemail-validate.sample:22:# Replace the TODO placeholders with appropriate c .git\hooks\sendemail-validate.sample:27: # TODO: Replace with appropriate checks (e.g. spe .git\hooks\sendemail-validate.sample:35: # TODO: Replace with appropriate checks for this .git\hooks\sendemail-validate.sample:41: # TODO: Replace with appropriate checks for the w .hidden-notes:1:TODO in a hidden file build\out.js:1:// TODO generated data\logo.bin: binary file matches (found "\0" byte around offset 4) debug.log:1:ERROR TODO in a log README.md:2:TODO write the docs scratch.txt:1:TODO untracked note src\app.js:1:// TODO(alice): handle timeout src\db.py:1:# TODO: use a connection pool

A claim worth testing: rg -uu (nothing ignored, hidden files searched, binary files still skipped) should behave like Chapter 2's grep -rnI --exclude-dir=.git. Run both, normalise the paths and sort, and compare:

diff <(grep -rnI --exclude-dir=.git TODO . | sed 's|^\./||' | sort) <(rg -uu -n -g '!.git' TODO | tr '\\' / | sort) && echo "identical output" echo "grep lines: $(grep -rnI --exclude-dir=.git TODO . | wc -l); rg lines: $(rg -uu -n -g '!.git' TODO | wc -l)\"
identical output ../_s.sh: line 5: unexpected EOF while looking for matching `"'

Choosing files

-t chooses a file type (rg knows a long list: rg --type-list), -g a glob, and a leading ! in a glob excludes:

rg -t js -n TODO
src\app.js:1:// TODO(alice): handle timeout
rg -g '*.py' -n TODO
src\db.py:1:# TODO: use a connection pool
rg -g '!README.md' -l TODO | sort
scratch.txt src\app.js src\db.py

The same options, and a few new ones

-i -w -x -F -v -c -l -o -A -B -C -e -f and -n mean what they mean in grep. A new one is smart case: -S ignores case if the pattern is all lower case, and is case-sensitive as soon as it has a capital. Plain rg is case-sensitive, like grep, and exits with 1 for no match:

rg todo -l; echo "status $?" rg -S todo -l | sort
status 1 README.md scratch.txt src\app.js src\db.py

-c counts matching lines per file, as grep does:

rg -c TODO | sort
README.md:1 scratch.txt:1 src\app.js:1 src\db.py:1

-z searches inside compressed files, as zgrep (Chapter 7) did:

printf 'a ERROR b\n' | gzip > x.gz rg -z ERROR x.gz
a ERROR b

A different regular expression language

rg's patterns are written in the Rust regex syntax: close to Chapter 3's extended form, with \d \w \s and lazy repeats, but with no look-around and no back-references. The error message points to the way round:

rg '(?<=TODO)\(' -n; echo status $?
rg: regex parse error: (?:(?<=TODO)\() ^^^^ error: look-around, including look-ahead and look-behind, is not supported Consider enabling PCRE2 with the --pcre2 flag, which can handle backreferences and look-around. status 2

-P switches to Perl-style patterns (Chapter 5), if rg was built with them; mine was:

rg -P '(?<=TODO)\(' -n
src\app.js:1:// TODO(alice): handle timeout

There is one more difference worth knowing. rg's \w understands Unicode, so it keeps accented letters in a word, where grep -P (Chapter 5) breaks them up. A file containing café résumé:

printf 'caf\xc3\xa9 r\xc3\xa9sum\xc3\xa9\n' > u.txt echo "rg:"; rg -o '\w+' u.txt echo "grep -P:"; grep -oP '\w+' u.txt
rg: café résumé grep -P: caf r sum

Where rg gets it wrong for you

rg honours .gitignore only inside a git repository, which is easy to forget when you copy files somewhere else. Outside one, the folder build is searched. --no-require-git makes rg obey the ignore file anyway, and so does running git init:

mkdir ../plain && cd ../plain printf 'build/\n' > .gitignore; mkdir build echo 'TODO a' > build/x.js; echo 'TODO b' > y.js echo "not a git repository:"; rg -n TODO | sort echo "with --no-require-git:"; rg --no-require-git -n TODO git init -q echo "after git init:"; rg -n TODO
not a git repository: build\x.js:1:TODO a y.js:1:TODO b with --no-require-git: y.js:1:TODO b after git init: y.js:1:TODO b

And when a search finds less than you expect, the first suspect is one of the skipping rules, which -u, -uu and -uuu will confirm. A missing file gives status 2 and a different message from grep's:

rg TODO nosuchfile; echo "status $?"
rg: nosuchfile: IO error for operation on nosuchfile: The system cannot find the file specified. (os error 2) status 2

The Silver Searcher (ag)

ag is a similar tool: from its documentation, it also searches recursively by default and respects ignore files. It is not installed here and I did not run it, so I will not describe its options. If you already use it, the ideas in this chapter apply: check what it skips by default.

How Much Faster?

A measurement, not a promise. The folder is this course's own content folder: 13,136 files, 12,470 of them tracked by git. Each command lists the files containing a word. First a word that occurs in 1,976 files, then one that occurs in none. Each ran three times; the numbers are seconds, as time measured them. The grep is the Windows build in Git Bash (3.0), the files are in a OneDrive folder, and I did not repeat this on Linux, so treat the figures as one data point from this machine:

(a script that ran grep -rlI, git grep -l and rg -l on the folder three times each; its output)
files in the folder: 13136; tracked by git: 12470 === pattern: copyCodeBlock matching files: grep -rlI: 1976 git grep -l: 1976 rg -l: 1976 grep -rlI seconds, three runs: 36.594 5.885 3.586 git grep -l seconds, three runs: 0.608 0.467 0.609 rg -l seconds, three runs: 0.553 0.460 0.512 === pattern: zzzqqqnotthere matching files: grep -rlI: 0 git grep -l: 0 rg -l: 0 grep -rlI seconds, three runs: 3.820 3.593 3.694 git grep -l seconds, three runs: 0.514 0.526 0.640 rg -l seconds, three runs: 0.511 0.539 0.588

Plain grep took about 3.6 seconds for each run once the files were cached, and 36.6 seconds on the very first run, when they had to be read from disk. git grep and rg took about half a second each, roughly seven times quicker than the warm grep, and all three found the same 1,976 files. The first run of a tool you have not used for a while is the slow one, which is why I ran each three times. On a project of a few dozen files, none of this matters.

grep on Windows

You have four ways to get grep-like searching on a Windows machine, and they are not the same:

OptionWhat it isNotes from this course
Git Basha Windows build of GNU grep, with gitwhat most examples ran on (grep 3.0); it hides carriage returns, so $ behaves differently from Linux (Chapter 3); it is the grep that was timed in the benchmark above (I did not time WSL's)
WSLa real Linux, with grep 3.11behaves like a Linux server; the right place to test a script that will run on one
PowerShell Select-Stringa PowerShell command, not grepbelow
findstra very old Windows command-line toolbelow

Select-String

Select-String does the job in PowerShell. It differs from grep in ways that catch people out. First, it is case-insensitive by default: here todo matches TODO. Its output is a match object, which prints in a familiar file:line:text form (PowerShell adds a blank line before the first result, left out here):

Select-String -Path src\app.js -Pattern todo
src\app.js:1:// TODO(alice): handle timeout

-CaseSensitive gives grep's behaviour:

Select-String -Path src\app.js -Pattern todo -CaseSensitive '(nothing above: -CaseSensitive found no lowercase todo)'
(nothing above: -CaseSensitive found no lowercase todo)

It does not recurse by itself. Feed it files from Get-ChildItem -Recurse. It knows nothing about .gitignore, so build\out.js is found:

Get-ChildItem -Recurse -File -Include *.js,*.py,*.md | Select-String TODO
build\out.js:1:// TODO generated src\app.js:1:// TODO(alice): handle timeout src\db.py:1:# TODO: use a connection pool README.md:2:TODO write the docs
grepSelect-String
-i (ignore case)the default; use -CaseSensitive to turn it off
-v-NotMatch
-F-SimpleMatch
-A N -B M / -C N-Context M,N (before, after)
-o-AllMatches and read .Matches.Value
-l-List (one match per file; then read the file name from the result)
-q-Quiet (returns True or False)
-rGet-ChildItem -Recurse -File | Select-String ...
Select-String -Path src\db.py -Pattern TODO -NotMatch
src\db.py:2:def connect(): src\db.py:3: # HACK: retry forever src\db.py:4: return None
Select-String -Path src\app.js -Pattern FIXME -Context 1,1
src\app.js:1:// TODO(alice): handle timeout > src\app.js:2:// FIXME validate input src\app.js:3:function start() { return legacyParse(input); }
Select-String -Path src\app.js -Pattern 'TODO\(\w+\)' -AllMatches | ForEach-Object { $_.Matches.Value } (Select-String -Path src\*.* -Pattern TODO -CaseSensitive).Count
TODO(alice) 2

The first shows -NotMatch (lines without TODO, with their line numbers); the second -Context 1,1, with > marking the match; the third pulls out just the matching text and counts the matching lines (case-sensitively).

There is no exit status for “no match”
grep says 1 when nothing matched. Select-String just returns nothing, and PowerShell's own status says all is well. To test for a match in a script use -Quiet, which returns True or False:
Select-String -Path src\app.js -Pattern NOPE -Quiet $r = Select-String -Path src\app.js -Pattern NOPE "no output object: $($null -eq $r); $? = $?"
False no output object: True; True = True

findstr

findstr is old, but it is present on every Windows machine and works from cmd. /n numbers lines, /i ignores case, /s recurses, /m lists file names only, and /c:"..." searches for a phrase with spaces. It sets an exit status like grep (0, or 1 for no match):

findstr /n /i /c:"use a connection" src\db.py echo status %errorlevel% findstr /s /m /i todo *.js *.py findstr /n NOPE src\db.py echo status when nothing found %errorlevel%
1:# TODO: use a connection pool status 0 build\out.js src\app.js src\db.py status when nothing found 1

Its regular-expression support is limited and I did not test it, so use grep, rg or Select-String when the pattern is more than words.

Which One Should I Use?

SituationUse
a script that has to run on any Linux, macOS or servergrep (-E, not -P)
searching a git project, every day, interactivelygit grep (no install) or rg
searching old versions of a projectgit grep PATTERN REVISION
a big folder, speed matters, and you can install a toolrg
files that git does not know about, or compressed logsgrep, zgrep or rg -z
a Windows machine with nothing but PowerShellSelect-String, and install Git Bash or WSL when you can
lookahead or back-references in a script for yourselfgrep -P or rg -P

A caution that applies to every tool in this chapter: their defaults hide files. A search that finds nothing might mean nothing is there, or that the tool skipped it. When it matters, check with the “do not skip anything” form (grep -r, rg -uuu, git grep --no-index).

Hands-On Exercises

Exercise 1

With git grep: list the tracked files that mention TODO, show which note existed in the first commit and is gone now, and find the line that has both TODO and alice. Then say which three kinds of file grep -r would have added.

📄 View solution
Exercise 2

Show that rg -uu and Chapter 2's grep -rnI --exclude-dir=.git find the same lines, then find where they differ if you drop -uu to -u, and say what each of the three -u levels turns off.

📄 View solution
Exercise 3

In PowerShell, write the Select-String versions of three greps from the repository (case-sensitive TODO in Python files; lines without TODO in db.py; the files that contain FIXME) and check each against grep's answer.

📄 View solution

Chapter 8 Quick Reference

  • git grep: tracked files, never .git; --untracked adds new files; --no-index works anywhere and includes ignored files; git grep PAT REV searches an old version; --cached the index; --and --or --not; -- '*.js' ':!*.md' pathspecs; status 128 outside a repo
  • rg: recursive by default; skips ignored, hidden and binary files; -u / -uu / -uuu turn off ignore, hidden, binary skipping; --files shows what it would search; -t, -g; -S smart case; -z compressed; -P for look-around (the default has none)
  • rg: output looks like grep's when piped (add -n); order is not stable (| sort); only obeys .gitignore in a git repo; Unicode-aware \w
  • Measured here: about 3.6 s for Git Bash grep against about 0.5 s for git grep and rg on 13,136 files (one machine, one afternoon)
  • Windows: Git Bash (grep 3.0, hides CR), WSL (Linux grep), Select-String (case-insensitive by default, no recursion, no ignore files, -Quiet for a yes or no), findstr (/n /i /s /m /c:, status 0 or 1)
  • Every tool's defaults hide files: when a search finds nothing, check with the version that skips nothing

Course Complete

That is all eight chapters. You can now search text with grep from a single file up to a whole project, write patterns in basic, extended and Perl form, use grep safely in scripts, and pick a better tool when grep is not the best one. What each chapter left you able to do:

ChapterYou can now…
1 · First Searchessearch for text; ignore case, invert, match whole words and whole lines; fixed strings; quote a pattern; read the exit status
2 · Searching Folderssearch a whole tree with -r; choose files and folders; deal with binary files, links and file names with spaces
3 · Basic and Extended Regexanchors, sets and classes, groups, counts and back-references; know which punctuation is an operator in each form
4 · Context and Output Controlshow lines around a match, line numbers and file names, count lines or matches, stay quiet, and count things with sort | uniq -c
5 · Perl Regex with -Puse lazy repeats, lookaround and \K; match across lines with -z; know where -P stops
6 · grep in Pipelines and Scriptsuse grep as a test in if; avoid the set -e, pipefail and blank-pattern traps; edit files safely after finding them
7 · Real-World Searchingread a log, hunt through a codebase without printing secrets, and audit configuration files
8 · Faster Alternativeschoose between grep, git grep, ripgrep and the Windows tools, with measured speed
Where next
The other tools in this family are on the site: SED — A Course for Beginners to change text, AWK to work with fields, and Learning Regular Expressions for the pattern language underneath all of them. The grep Cheat Sheet in the Linux part of the Sidebar is the one-page reminder of everything in Chapters 1 to 5.