learning-website-framework1-8 Exercise 3: Find Secrets Before You Publish ======================================================================= Content that is published, committed or synced can contain real credentials (a database password pasted into a lesson, a key in a script). Write two scanners and run them on the real folders: 1. a PATTERN scan: finds things that look like secrets (NAME = value); 2. a KNOWN-SECRET scan: finds exact strings you already know are secret. Both must show nothing of the values they find. Save as scan_secrets.py: import os, re, sys from collections import Counter TEXT_EXT = {".html", ".txt", ".md", ".py", ".sh", ".js", ".ts", ".json", ".env", ".yaml", ".yml", ".cfg", ".conf", ".php"} SKIP_DIRS = {".git", "node_modules", "__pycache__", "pdfs"} # NAME = value / NAME: value (quotes optional); the value stops at a quote, space or "<" ASSIGN = re.compile(r"""(?ix) \b(?P[A-Za-z_]*(?:pass(?:word|wd)?|secret|token|api[_-]?key)) \s*[=:]\s*['"]?(?P[A-Za-z0-9][^\s'"<>]{7,})""") PLACEHOLDER = re.compile(r"(?i)(your|example|changeme|placeholder|xxx|\*{3}|<|\$\{|\$\(|replace|secret_?key|password|token)") def looks_real(value): """Needs 3 of 4 character classes and must not look like a placeholder.""" if PLACEHOLDER.search(value): return False classes = sum([bool(re.search(r"[a-z]", value)), bool(re.search(r"[A-Z]", value)), bool(re.search(r"\d", value)), bool(re.search(r"[^A-Za-z0-9]", value))]) return classes >= 3 def mask(value): return "*" * len(value) # show nothing of the value, not even its first character def scan_file(path): hits = [] try: text = open(path, encoding="utf-8", errors="ignore").read() except OSError: return hits for n, line in enumerate(text.splitlines(), 1): for m in ASSIGN.finditer(line): if looks_real(m.group("val")): hits.append((n, m.group("key"), len(m.group("val")), mask(m.group("val")))) return hits def scan(root): results = [] for dirpath, dirnames, filenames in os.walk(root): dirnames[:] = [d for d in dirnames if d not in SKIP_DIRS] for name in filenames: if os.path.splitext(name)[1].lower() in TEXT_EXT: p = os.path.join(dirpath, name) for hit in scan_file(p): results.append((os.path.relpath(p, root).replace(os.sep, "/"), *hit)) return results if __name__ == "__main__": for root in sys.argv[1:]: res = scan(root) files = {r[0] for r in res} print(f"{root}") print(f" {len(res)} possible secrets in {len(files)} files") for rel, n, key, length, masked in res[:12]: print(f" {rel}:{n} {key} ({length} chars) {masked}") Save as scan_known.py (the list of known secrets lives in a file OUTSIDE the repository, one per line): import os, sys SKIP_DIRS = {".git", "node_modules", "__pycache__"} TEXT_EXT = {".html", ".txt", ".md", ".py", ".sh", ".js", ".ts", ".json", ".env", ".yaml", ".yml", ".cfg", ".conf", ".php", ".css", ".xml"} def load_known(path): """One secret per line, kept OUTSIDE the repository. Blank lines and # comments ignored.""" with open(path, encoding="utf-8") as fh: return [ln.rstrip("\n") for ln in fh if ln.strip() and not ln.startswith("#")] def scan_known(root, secrets): """Return {relative file: number of secrets found}; never returns the values.""" found = {} for dirpath, dirnames, filenames in os.walk(root): dirnames[:] = [d for d in dirnames if d not in SKIP_DIRS] for name in filenames: if os.path.splitext(name)[1].lower() not in TEXT_EXT: continue # skip PDFs, images and other binary files p = os.path.join(dirpath, name) try: text = open(p, encoding="utf-8", errors="ignore").read() except OSError: continue n = sum(text.count(s) for s in secrets) if n: found[os.path.relpath(p, root).replace(os.sep, "/")] = n return found if __name__ == "__main__": known_file, *roots = sys.argv[1:] secrets = load_known(known_file) print(f"checking {len(secrets)} known secret(s); values are never printed") for root in roots: hits = scan_known(root, secrets) print(f"{root}: {len(hits)} file(s)") for rel, n in sorted(hits.items()): print(f" {rel} ({n} occurrence{'s' if n != 1 else ''})") Run them: python scan_secrets.py "/claude-projects/website-content/content" python scan_known.py known_secrets.txt "/claude-projects/website-content/content" "/dist" Output of the pattern scan, first lines (checked by running it): /claude-projects/website-content/content 66 possible secrets in 46 files linux/shell-and-scripting/bash-scripting/advanced/bash_advanced_07.html:638 nAPI_KEY (11 chars) *********** linux/shell-and-scripting/sed/sed_lesson_09.html:518 password (10 chars) ********** programming/apis-and-services/graphql/graphql1-8.html:294 token (11 chars) *********** programming/compiled-and-systems-languages/go/advanced/solutions/go3-4_challenge3.txt:63 PASS (30 chars) ****************************** programming/compiled-and-systems-languages/go/advanced/solutions/go3-4_challenge3.txt:65 PASS (29 chars) ***************************** programming/developer-tooling/git/foundations/git_github_foundations_6.html:260 API_KEY (28 chars) **************************** programming/web-backend-frameworks/django/intermediate-advanced/django2-1.html:327 password (17 chars) ***************** programming/web-backend-frameworks/django/intermediate-advanced/solutions/django2-1_challenge1.txt:14 password (17 chars) ***************** projects/building-a-web-framework/building_a_web_framework_1_10.html:232 password (10 chars) ********** projects/catalogue-react-firebase/personal_catalogue_react_firebase_1_9.html:166 apiKey (38 chars) ************************************** projects/catalogue-react-firebase/personal_catalogue_react_firebase_1_9.html:177 VITE_FIREBASE_API_KEY (12 chars) ************ projects/catalogue-react-firebase/personal_catalogue_react_firebase_1_9.html:182 VITE_FIREBASE_API_KEY (13 chars) ************* Output of the known-secret scan (checked by running it; one real database password was loaded from a local script into known_secrets.txt, and no value is printed): checking 1 known secret(s); values are never printed /claude-projects/website-content/content: 5 file(s) /astro-site/dist: 6 file(s) What the scans show ------------------- - The pattern scan is a heuristic. It reports 66 candidates in 46 files, almost all of them example passwords and placeholders in lessons, and it MISSED 3 of the 5 files that really contain the known password (it caught 2). Use it to find candidates, not to prove a folder is clean. - The known-secret scan found that password in 5 content files (4 chapters and a solution file), in the 5 matching built pages, and in one more file in the build. (The file names are hidden in the output above on purpose; they are listed in the private review file, not published.) These are real findings, not fixed here. - Findings in a published page are PUBLIC if that page has been deployed. If a real credential turns up ----------------------------- 1. Change (rotate) the credential first. Anything that has been published, even briefly, must be treated as exposed. Removing the text does not undo that. 2. Replace it in the content with an obvious placeholder (for example your-db-password-here), then rebuild and redeploy. 3. If it was committed to git, rewriting the history removes it from the repository but not from copies already cloned or pushed; rotation is what makes it safe. 4. Make the known-secret scan a step that FAILS the build, so it cannot happen again. WHY THIS WORKS AS AN ANSWER --------------------------- Pattern scans find the unknown and exact-match scans find the known; neither alone is enough, as the result shows. Keeping the secret list outside the repository, never printing values and treating rotation as the real fix are the habits that matter.