PDFs, Solutions and Assets

Learning Website with Next.js

Chapter 8 ยท PDFs, Solutions & Assets

A page is not the only thing a visitor downloads. Courses have PDFs, chapters link to solution files, and 1,923 pages have a code block with a Copy button. This chapter sorts out how each is handled in Next.js, keeps the addresses that old pages already link to, and finds out whether the image and font tools that Next.js offers have anything to do here. Each part was checked on the real files, and the Copy button in a real browser.

Run for real, on 8,431 files and 4,419 pages
63 tests pass. The assets of every site were counted, the languages site's were copied and served, the headers were read with curl, and a headless Chrome clicked the Copy button through 12 checks. Nothing was run behind Apache, and the real clipboard was never written (that would overwrite what you had copied).

Three Kinds of File

KindExamplesHow it is handled
Pages4,419 content filesMade at build time (Chapter 4)
AssetsCourse PDFs, solution .txt files (8,431)Copied into the app's public/ folder at the same relative path
The site's own filesStylesheets, scripts (including the Copy button's)Compiled and hashed by Next

Assets: Same Address, Copied In

A page links to a PDF or a solution at the same relative path as in the content folder: /hungary/hungarian-basic-3/solutions/ex1.txt. Those addresses must keep working after the split. Next serves everything in an app's public/ folder as it is, at the same path, so each site's build copies its own PDFs and solutions there:

  • A site gets only its own files. The site is decided from the path (the same function as for pages), and a file that belongs to no site is reported and never copied.
  • Only what changed is copied (missing, or a different size or time), and the copy keeps the original's modified time, so a second run copies nothing.
  • --dry-run shows what would happen; --prune removes only files that a previous run copied and that no longer exist.
  • PDFs are build artefacts. The Python builders make them, git ignores them (*.pdf), and this build only copies the ones that exist. On a fresh checkout there are none until the builders have run.
Real runResult
Languages, first copy236 files, 186 MiB, 1.6 s
The same again0 copied, 236 unchanged, 0.3 s
Every site, dry runLanguages 236, web development 1,016, programming 3,658, systems 1,419, AI 317, humanities 889, life skills 412, creative 484: 8,431 files, none skipped (about 640 MiB of copies across eight apps)

Serving Them

Request (curl, running app)Result
A solution .txt200, text/plain; charset=UTF-8, ETag, Last-Modified; a repeat with the ETag gives 304
A PDF200, application/pdf, Accept-Ranges: bytes; a request for bytes 0–99 gives 206 Partial Content
A missing file under a solutions folder404
/solutions/../../../package.json and the %2e%2e form404, 404

One header is added to the two folders: X-Content-Type-Options: nosniff, so a browser never guesses a file's type. To see why, a file containing <script>alert(1)</script><h1>hi</h1> was put in a solutions folder (and removed afterwards). It is served as text/plain, and Chrome showed it as text with the tags escaped: nothing ran.

Three mistakes, and how to avoid each
  1. A bookkeeping file inside public/. My list of copied files (needed for --prune) was kept in the target folder, and public/ is served to visitors. A request for it was answered, with an error (500), which is its own bad sign. It is now kept next to the folder, and a test checks that only content is in public/. Whatever a folder serves, nothing in it should be for you only.
  2. A wrong comment. I wrote that Next sends no validators for files in public/, and added a one-hour cache rule. I had filtered the header list too tightly to see them: Next does send ETag and Last-Modified, and its default max-age=0 means “ask again each time, cheaply”. The rule and the comment were removed. Look at all the headers before explaining them.
  3. A test that assumed a language. My “only content is in public” test expected one folder and found two, because its own sample files included a French one and French belongs to the languages site. Work out what a test expects from the same rule the code uses.

Downloads on Course Pages, and Dead Links

A course folder's page now lists its PDFs, and its solution files folded away in a <details>, taken from what was really copied, so a PDF that does not exist cannot be linked. /hungary/hungarian-basic-3 shows a Downloads heading with its course PDF and chapter PDFs; /hungary, which has none of its own, shows no heading.

A second check goes the other way: every link to a .pdf or .txt in every page, against the files the builds would copy. 11,188 links, 8,431 files, 178 links to a file that does not exist: 174 on the web development site (solution files for the scripting and back-end courses), 3 on creative, 1 on systems. It is the same number the Django project found, reached independently. It was not made to fail the build, because 178 known problems would stop every build until someone decides what to do about them.

Images and Fonts: Is There Anything to Optimise?

Next.js has next/image (resizes and converts images) and next/font (loads fonts without a layout jump). Whether to use them depends on what the pages really contain, so the fragments were surveyed:

Found in the 4,419 pagesCount
<img> tags0
Image or font files in the content folder0
Pages with @font-face5 (all web development)
Pages loading Google Fonts3 (2 web development, 1 systems)
Pages with url() in their CSS40
Pages with a Copy button1,923 (none on the languages site)

So next/image has nothing to work on: no page uses an image, and the old site's pictures live in its resources/ folder (code samples and exercises), not in the content. The site's fonts are the system fonts (Chapter 5), and the few pages with their own fonts bring them themselves, so next/font would not replace those either. Neither is adopted. Revisit this when the first image is added to a page, and use next/image for it then.

The Copy Button as a Client Component

On 1,923 pages a code block is written as <button class="copy-btn" onclick="copyCodeBlock(this)">. The page body is inserted as text, so a script inside it would never define the function (Chapter 4). It is defined once, for the whole site, by a small client component in the layout. The button copies the code block right after it, using the browser's clipboard when it is available (HTTPS or localhost), the old hidden-text-box command when it is not, and saying “Copy failed” if both fail.

Headless Chrome, on the test benchResult
The page defines copyCodeBlock; the button starts as “Copy”pass
Modern clipboard: asked to copy exactly the code block's text; button says “Copied!”; back to “Copy” after 1.5 spass, pass, pass
The clipboard refuses: the old command is used, on the right text; “Copied!”pass, pass
No clipboard at all: the old command is usedpass
Nothing works: the button says “Copy failed”; no code block after the button: “Copy failed”, no crashpass, pass
After a link click: the function still exists and copies the right textpass, pass
What this test does not do, and whether it can fail
It never writes the real system clipboard, on purpose: that would overwrite whatever you had copied. The browser's clipboard call and the old copy command are replaced by recorders, so the test sees which text is asked for, which route is taken and what the button says. The last step, pasting, has not been seen: please click a Copy button on a real page once. All 12 passed first time, which proves little, so two deliberate breakages were rebuilt and run: a button that always claims “Copied!” gave 11 passed, 1 failed; copying the element before the button gave 6 passed, 6 failed. Be wary of a test until you have seen it fail for the right reason.
What was not verified
Apache is not in front, so none of the serving was checked behind a proxy (in production Apache can serve these folders directly: Chapter 12). The PDFs themselves were not rebuilt. The Copy button was tested in Chrome only, and never against the other 1,922 pages' real code blocks, only the test bench's.

Hands-On Exercises

Exercise 1

Copy each site's PDFs and solutions into its app at the same relative paths, copying only what changed and pruning only what you made. Serve them and check the types, ranges, cache validators and path tricks with curl. Record the mistakes you make.

๐Ÿ“„ View solution
Exercise 2

List a course's downloads on its page from the files that really exist, check every PDF and solution link in every page, and survey the pages' images and fonts to decide whether next/image and next/font are worth using.

๐Ÿ“„ View solution
Exercise 3

Make the Copy button work on every page with one client component, and test every route it can take in a real headless Chrome without touching the real clipboard. Show the test fails when the button lies.

๐Ÿ“„ View solution

Chapter 8 Quick Reference

  • Assets keep their same relative path: a site's build copies its own PDFs and solutions into the app's public/
  • Only changed files are copied (size and time); --prune removes only what an earlier run copied; the list of copied files is kept next to public/, never in it
  • 8,431 files in all (about 640 MiB of copies); languages 236 files, 186 MiB, 1.6 s the first time, 0.3 s after
  • PDFs are built by the Python builders and ignored by git: the site build copies the ones that exist
  • Next sends ETag and Last-Modified for public/ files (304 on a repeat), range requests give 206; add nosniff for /pdfs/ and /solutions/
  • A course page lists its PDFs and (folded) its solutions, from the files that exist
  • 11,188 links to PDFs and solutions checked: 178 point at missing files (174 web development), the same number as Django
  • 0 images and 0 font files in the content: next/image and next/font not adopted
  • Copy button: one client component in the layout defines copyCodeBlock; modern clipboard, then the old command, then "Copy failed"; 12 checks pass, and the test fails when the button lies
  • Not done: Apache in front, rebuilding the PDFs, a real paste, other browsers