Capstone: Designing and Critiquing a Real Experiment
Scientific Methodology
Chapter 10 · Capstone: Designing and Critiquing a Real Experiment
This capstone works in two directions. First, it designs a genuinely falsifiable experiment from scratch, applying every tool from Chapters 1-9 in the order a real researcher would actually need them. Second, it turns that same toolkit backward — onto the real Santa Clara antibody study from Chapter 8 — to ask a concrete question: which of these nine chapters' tools, applied before publication, would have caught the flaw that only became visible afterward?
Part 1: Designing a Falsifiable Experiment From Scratch
Working hypothesis: does studying with background instrumental music change next-day recall test scores compared with studying in silence?
The hypothesis is framed in a form that could genuinely fail: "instrumental background music produces no significant difference in next-day recall scores compared with silence." A real result — in either direction — can disconfirm it, satisfying Popper's own falsifiability criterion.
This single study tests one specific hypothesis. It is not, by itself, a theory — it would need to be one small, repeated piece of evidence feeding into a broader theory (e.g. cognitive load theory) only after independent replication, never treated as confirming that broader theory on its own.
Independent variable: music vs. silence during study. Dependent variable: next-day recall test score. Controlled: identical study material, identical study duration, identical test. Control group: the silence condition. Randomization: participants randomly assigned to condition, not self-selected.
Participants can't be blinded to their own condition — they know whether they're hearing music. An honest limitation, not a flaw to hide. But the researcher scoring each recall test can be blinded to which condition each participant was in, preventing the same kind of expectation-driven scoring bias Chapter 4's placebo material covered.
Sample size is set in advance via a real power calculation, not by "running until significant" — the exact flexible-analysis pattern behind the Open Science Collaboration's own 36% replication rate. The full analysis plan is fixed before data collection begins.
What result would count against the hypothesis is written down in advance — and if the result comes back null or unexpected, no new, untested sub-group or post hoc statistical construct will be invented to rescue it, the exact ad hoc pattern that discredited the real Mars effect.
The introduction and methodology are submitted for real peer review before data collection — the Center for Open Science's own Registered Reports format, which measurably cut psychology's suspicious 96% "positive result" rate down to 44% by removing the incentive to chase a significant finding.
Because the study is a Registered Report, publication doesn't depend on the result — closing off the file-drawer problem before it can happen. In advance, the plan also names what would trigger a real correction (a fixable error), an expression of concern (a serious, unresolved doubt), or a full retraction (a genuinely invalidated result) if a flaw were discovered after the fact.
Part 2: Critiquing a Real Experiment After the Fact
Chapter 8 covered the real Santa Clara antibody study (April 2020) in detail — a preprint claiming COVID-19 infections were undercounted 50-85x, criticized for not properly accounting for its antibody test's own real specificity uncertainty, and eventually revised and published in 2021 at a lower estimate. Applying this capstone's own nine-chapter toolkit backward asks: how much of that criticism could have been caught before the preprint ever reached the public?
| Tool | Applied to Santa Clara in advance |
|---|---|
| Ch.1 Falsifiability | The core claim was genuinely falsifiable — this was never the problem |
| Ch.3 Variable design | The test's own real false-positive rate wasn't fully incorporated into the reported confidence — a design-stage gap, not a data-collection error |
| Ch.6 Pre-registered analysis | The gap this chapter exists to close: no pre-registered plan specifying in advance how test-specificity uncertainty would be handled |
| Ch.7 No ad hoc rescue needed | To the study's credit, no post hoc rescue was invented once flaws were raised — the estimate was genuinely revised, not defended by new untested constructs |
| Ch.8 Registered Report | Not used — a regular preprint reached massive public and policy attention before independent statisticians could review the specificity math |
| Ch.9 Self-correction | Not retracted — revised and formally published in 2021 at a lower, still-elevated estimate, closer to a large-scale correction than a full retraction |
Had the Santa Clara design gone through a Registered Report process (Ch.8) — requiring the antibody test's own real specificity uncertainty to be pre-specified and reviewed in the analysis plan (Ch.6) before data collection — the exact criticism that eventually forced a public revision could plausibly have been raised and resolved at the design stage, before the preprint ever reached policy-relevant public attention. The tools this course covers aren't only useful in hindsight; several of them exist specifically to catch a problem like this one before it happens.
Chapter Attribution
| Capstone step | Chapter |
|---|---|
| Falsifiable hypothesis | Chapter 1 |
| Hypothesis vs. theory | Chapter 2 |
| Variables, control group, randomization | Chapter 3 |
| Blinding | Chapter 4 |
| Reflecting on paradigm resistance to a null result | Chapter 5 |
| Pre-specified sample size and analysis plan | Chapter 6 |
| Pre-specified falsification criteria, no ad hoc rescue | Chapter 7 |
| Registered Report submission | Chapter 8 |
| Publish-regardless-of-outcome, correction/EoC/retraction plan | Chapter 9 |
Course Complete — Study Methodologies Subject
- Research Methodology (10 chapters) — sourcing, credibility, and study design fundamentals
- Critical Thinking (10 chapters) — everyday reasoning, fallacies, and cognitive biases
- Scientific Methodology (10 chapters) — demarcation, experimental design, paradigm shifts, reproducibility, and self-correction
- The Study Methodologies Subject is now complete: 30 chapters across three courses