Capstone: Designing and Critiquing a Real Experiment

Scientific Methodology

Chapter 10 · Capstone: Designing and Critiquing a Real Experiment

This capstone works in two directions. First, it designs a genuinely falsifiable experiment from scratch, applying every tool from Chapters 1-9 in the order a real researcher would actually need them. Second, it turns that same toolkit backward — onto the real Santa Clara antibody study from Chapter 8 — to ask a concrete question: which of these nine chapters' tools, applied before publication, would have caught the flaw that only became visible afterward?

Part 1: Designing a Falsifiable Experiment From Scratch

Working hypothesis: does studying with background instrumental music change next-day recall test scores compared with studying in silence?

Step 1 — Ch.1: State It Falsifiably

The hypothesis is framed in a form that could genuinely fail: "instrumental background music produces no significant difference in next-day recall scores compared with silence." A real result — in either direction — can disconfirm it, satisfying Popper's own falsifiability criterion.

Step 2 — Ch.2: Distinguish Hypothesis From Theory

This single study tests one specific hypothesis. It is not, by itself, a theory — it would need to be one small, repeated piece of evidence feeding into a broader theory (e.g. cognitive load theory) only after independent replication, never treated as confirming that broader theory on its own.

Step 3 — Ch.3: Define the Real Variables

Independent variable: music vs. silence during study. Dependent variable: next-day recall test score. Controlled: identical study material, identical study duration, identical test. Control group: the silence condition. Randomization: participants randomly assigned to condition, not self-selected.

Step 4 — Ch.4: Blind What Can Be Blinded

Participants can't be blinded to their own condition — they know whether they're hearing music. An honest limitation, not a flaw to hide. But the researcher scoring each recall test can be blinded to which condition each participant was in, preventing the same kind of expectation-driven scoring bias Chapter 4's placebo material covered.

Step 5 — Ch.6: Design for Reproducibility Before Running It

Sample size is set in advance via a real power calculation, not by "running until significant" — the exact flexible-analysis pattern behind the Open Science Collaboration's own 36% replication rate. The full analysis plan is fixed before data collection begins.

Step 6 — Ch.7: Pre-Specify What Would Count as Failure

What result would count against the hypothesis is written down in advance — and if the result comes back null or unexpected, no new, untested sub-group or post hoc statistical construct will be invented to rescue it, the exact ad hoc pattern that discredited the real Mars effect.

Step 7 — Ch.8: Submit as a Registered Report

The introduction and methodology are submitted for real peer review before data collection — the Center for Open Science's own Registered Reports format, which measurably cut psychology's suspicious 96% "positive result" rate down to 44% by removing the incentive to chase a significant finding.

Step 8 — Ch.9: Commit to Publishing Regardless of Outcome

Because the study is a Registered Report, publication doesn't depend on the result — closing off the file-drawer problem before it can happen. In advance, the plan also names what would trigger a real correction (a fixable error), an expression of concern (a serious, unresolved doubt), or a full retraction (a genuinely invalidated result) if a flaw were discovered after the fact.

Part 2: Critiquing a Real Experiment After the Fact

Chapter 8 covered the real Santa Clara antibody study (April 2020) in detail — a preprint claiming COVID-19 infections were undercounted 50-85x, criticized for not properly accounting for its antibody test's own real specificity uncertainty, and eventually revised and published in 2021 at a lower estimate. Applying this capstone's own nine-chapter toolkit backward asks: how much of that criticism could have been caught before the preprint ever reached the public?

ToolApplied to Santa Clara in advance
Ch.1 FalsifiabilityThe core claim was genuinely falsifiable — this was never the problem
Ch.3 Variable designThe test's own real false-positive rate wasn't fully incorporated into the reported confidence — a design-stage gap, not a data-collection error
Ch.6 Pre-registered analysisThe gap this chapter exists to close: no pre-registered plan specifying in advance how test-specificity uncertainty would be handled
Ch.7 No ad hoc rescue neededTo the study's credit, no post hoc rescue was invented once flaws were raised — the estimate was genuinely revised, not defended by new untested constructs
Ch.8 Registered ReportNot used — a regular preprint reached massive public and policy attention before independent statisticians could review the specificity math
Ch.9 Self-correctionNot retracted — revised and formally published in 2021 at a lower, still-elevated estimate, closer to a large-scale correction than a full retraction
The Real, Concrete Finding

Had the Santa Clara design gone through a Registered Report process (Ch.8) — requiring the antibody test's own real specificity uncertainty to be pre-specified and reviewed in the analysis plan (Ch.6) before data collection — the exact criticism that eventually forced a public revision could plausibly have been raised and resolved at the design stage, before the preprint ever reached policy-relevant public attention. The tools this course covers aren't only useful in hindsight; several of them exist specifically to catch a problem like this one before it happens.

Chapter Attribution

Capstone stepChapter
Falsifiable hypothesisChapter 1
Hypothesis vs. theoryChapter 2
Variables, control group, randomizationChapter 3
BlindingChapter 4
Reflecting on paradigm resistance to a null resultChapter 5
Pre-specified sample size and analysis planChapter 6
Pre-specified falsification criteria, no ad hoc rescueChapter 7
Registered Report submissionChapter 8
Publish-regardless-of-outcome, correction/EoC/retraction planChapter 9
What This Course Doesn't Cover
This course covers the philosophy and mechanics of scientific methodology itself — demarcation, experimental design, paradigm change, reproducibility, and self-correction. It doesn't cover statistical analysis techniques in depth (see this Subject's own sibling course, Research Methodology, for source evaluation and study-design basics), formal logic and argument evaluation (see Critical Thinking), or the mathematics of probability and inference (see the site's own Maths for Programmers Subject). This course is the philosophy-and-practice layer underneath both sibling Study Methodologies courses, not a replacement for either.

Course Complete — Study Methodologies Subject

  • Research Methodology (10 chapters) — sourcing, credibility, and study design fundamentals
  • Critical Thinking (10 chapters) — everyday reasoning, fallacies, and cognitive biases
  • Scientific Methodology (10 chapters) — demarcation, experimental design, paradigm shifts, reproducibility, and self-correction
  • The Study Methodologies Subject is now complete: 30 chapters across three courses