The Languages Site

Learning Website with Django

Chapter 7 · The Languages Site

So far everything has been shared by all eight sites. This chapter builds the first thing that belongs to just one: the languages site, the area chosen to move first (Learning Website: Framework & Architecture 10). It needs a few things the others do not: it must know which language is which, group its courses the way a learner thinks, and mark the language of every line being taught. The lesson is as much about where that code goes as about what it does: a small app of its own, plugged into the shared pipeline, with nothing language-specific added to the shared apps.

Run for real, on Django 6.1
The project has 152 tests (all passing). The changes were run on the real content: 561 languages pages, and the results were checked against counts taken from those pages before the code was written.

Count First

Before writing the markup code, the classes the lesson fragments really use were counted. Each line in the language being taught sits in an element with a class named after the language, and the English line has the class en:

ClassLinesPagesHTML language code
de (German)1,28548de
hu (Hungarian)1,27463hu
jp (Japanese)68735ja, not jp
fr (French)36315fr

The lesson wrappers are .jp-lesson (95 pages), .hu-lesson (67), .de-lesson (48) and .fr-lesson (15), plus four more used by the Japanese culture courses, which have no colour of their own yet.

Why Mark the Language?

  • Pronunciation. A screen reader switches to the right voice for lang="hu" text. Without it Hungarian is read with English rules.
  • Glyph shapes. The same kanji are drawn with different shapes for Chinese and Japanese. lang="ja" tells the browser to use the Japanese forms.
  • Spelling and hyphenation use the right language.
<span class="jp">おはようございます。</span> becomes <span lang="ja" class="jp">おはようございます。</span>

The pass is small, and the care is in the edges. It reads whole class names (so hu matches but hu-lesson does not), leaves a tag that already has a lang attribute alone, copes with attributes spread over several lines, ignores text that merely mentions a class (an escaped example or a CSS rule), and gives the same result when run twice. Each of those is a test.

Plugging It In Without Touching the Shared Apps

The shared importer should not know about Hungarian. Instead, apps/content/fragments.py keeps a small registry, and the languages app adds its function to it when Django starts:

# apps/site_languages/apps.py def ready(self): from apps.content.fragments import register_transform from .markup import add_language_attributes register_transform("languages", add_language_attributes)

The importer calls prepare_fragment(raw, path, site), which runs the shared steps and then that site's registered functions. Registering twice does not run a function twice, and a page of any other site is untouched. The real check: after importing, there are 0 of these attributes on any other site, and on the languages site the counts match the table above exactly (1,285, 1,274, 687 and 363).

A Gotcha: When the Pipeline Changes but the Files Do Not

Chapter 4's importer skips a file whose hash has not changed. But the hash says whether the file changed, not whether the way it is processed did. After adding the markup pass, every page would still say “unchanged” and the new attributes would only reach pages someone happened to edit.

PIPELINE_VERSION = "2" # raise it whenever a processing step changes digest = hashlib.sha256(PIPELINE_VERSION.encode() + b"\0" + data).hexdigest()

With the version inside the hash, raising it re-imports every page once and then settles. On the real content: 4,403 updated, then 4,406 unchanged on the next run (the extra three are new chapters added in the meantime). A test does the same on a temporary folder.

The Front Page

The languages site gets a front page of its own, built from the courses in the database. The courses are grouped by the rule recorded when the plan for six courses in two volumes was made:

GroupWhat goes in itReal example
Reading and writingAlphabet, hiragana, katakana, kanji courses, in that orderHungarian Alphabet 1; Hiragana 1 to 3; Katakana 1 and 2
Survival <Language>Basic Conversation courses 1 to 3Survival German: courses 1, 2, 3
Everyday <Language>Basic Conversation courses 4 to 6Everyday German: course 4
CultureCourses about the countryJapanese Music; Manga & Anime; five in all
Lessons and referenceStandalone lessons, not numbered courses12 Hungarian pages, 3 French, 300 Japanese

The groups are produced by one small function with tests, so the rule is written down once, in code. The result on the real data: French 1 course; German 4 (Survival German and Everyday German); Hungarian 4 (Reading and writing, Survival Hungarian); Japanese 11 (Reading and writing, Survival Japanese, Culture). A language with no numbered courses still has its card; a lessons link only appears where lessons exist.

Colours, and a fix for the Japanese red

Each language's card sets two variables: --lang-accent for the border and --lang-text for the heading. They are equal for French, German and Hungarian. Japanese is the exception that Learning Website: Framework & Architecture 4 predicted: its red, #d64550, scores only 3.86 against the dark surface, under the 4.5 that WCAG AA asks for normal text. The heading therefore uses a lighter red of the same hue, #e8707a, which scores 5.63. One test checks every language's text colour, and another checks that the original red is below 4.5, so nobody “simplifies” the variant away later.

Answers Without a Script

The “Try It Yourself” answers are in a native <details> element. It opens and closes with no JavaScript, so it works in the stored pages exactly as written. 196 of the 561 languages pages use it.

What Is Not Done Here

  • The Japanese culture courses use their own wrapper classes (jfilmtv1-lesson, jlit1-lesson and so on), which have no colour in the shared tokens, so they appear in the site's accent. Adding four lines to tokens.css would fix it.
  • PDFs and solution files are the next chapter.
  • The front page was looked at in a desktop-width screenshot only.

Hands-On Exercises

Exercise 1

Count the language classes in the real pages, then build the languages app: a config of languages, a markup pass that adds lang attributes (whole class names only, idempotent), and a registry in the shared pipeline so the app plugs in without changing the shared code.

📄 View solution
Exercise 2

Group the courses (writing, survival, everyday, culture), build the front page with a card per language, give Japanese a readable text colour, and let a site name a front page of its own. Test the grouping, the colours and the front page.

📄 View solution
Exercise 3

Add a pipeline version to the hash so a change in processing re-imports every page once. Import the real content, and check that the stored language attributes match the counts you took before writing the code, that no other site changed, and that the front page groups the real courses correctly.

📄 View solution

Chapter 7 Quick Reference

  • Site-specific code lives in its own app (apps/site_languages) and plugs in through register_transform(site, function) in ready()
  • Count the real markup before coding: 1,285 de, 1,274 hu, 687 jp, 363 fr lines
  • lang attributes fix pronunciation, Japanese glyph shapes and spelling; Japanese is ja, not jp
  • The markup pass: whole class names only, skips existing lang, multi-line tags, idempotent
  • PIPELINE_VERSION inside the file hash: processing changes re-import every page once (4,403 updated, then unchanged)
  • Groups: Reading and writing, Survival (courses 1 to 3), Everyday (4 to 6), Culture; Lessons and reference linked where they exist
  • Per-language card colours: --lang-accent and --lang-text; Japanese text uses #e8707a (5.63) because #d64550 is 3.86
  • Answers use a native <details>: no script needed (196 pages)
  • Not yet: colours for the Japanese culture courses, PDFs and solution files