Peer Review, Preprints & the Open Science Movement
Scientific Methodology
Chapter 8 · Peer Review, Preprints & the Open Science Movement
The sibling course Research Methodology already covers how traditional peer review works, how to assess source credibility, and real cases of a broken peer-review pipeline. This chapter picks up a different thread: real, current reforms — preprints and open science — built directly in response to problems like the reproducibility crisis from Chapter 6.
Preprints: Publishing Before Formal Review
A preprint is a completed research paper posted publicly before undergoing formal peer review — trading the months-long review timeline for immediate, open visibility.
Physicist Paul Ginsparg launched arXiv at Los Alamos National Laboratory on August 14, 1991, as a simple email-based server for high-energy physics theory papers — expected to handle roughly 100 submissions a year, it received about 400 in its first six months alone. Ginsparg moved it to Cornell University in 2001, where it's hosted today. It has since expanded well beyond physics into math, computer science, and more. By mid-2026 arXiv had passed 3 million hosted articles, with monthly submissions repeatedly setting new records — over 30,000 papers in a single month by 2026, up from roughly 120,000/year in 2017.
bioRxiv (2013) extended the same model to biology; medRxiv (2019) extended it to medicine — with one real, deliberate difference: medRxiv applies stricter screening and more prominent disclaimers than bioRxiv. The reason is direct real-world stakes — an unvalidated biology finding rarely changes anyone's behavior overnight, but an unvalidated clinical claim can influence real treatment decisions and public policy the moment it reaches the news. The case below shows exactly that risk playing out.
A Real Case: The Santa Clara Antibody Study
In April 2020, a Stanford-affiliated team — including Eran Bendavid, Jay Bhattacharya, and John Ioannidis (whose 2005 paper Chapter 6 already covered) — posted a preprint to medRxiv testing 3,330 Santa Clara County residents for COVID-19 antibodies.
50 of 3,330 tests (1.5%) came back positive. The preprint's headline conclusion: the true infection rate was 50–85 times higher than confirmed case counts — implying COVID-19 was far more widespread, and far less deadly, than official numbers suggested. Released amid an active lockdown-policy debate, it drew immediate, massive public and media attention.
The study was not retracted. It went through substantial revision and, after formal peer review, was published in 2021 in the International Journal of Epidemiology — with the estimate revised down to a seroprevalence of 2.8%, translating to roughly a 54-fold undercount, closer to (but still below) the original claim. This is a genuinely nuanced real outcome: peer review functioned as intended, just applied to a claim that had already reached wide public attention before that scrutiny happened — exactly the structural risk a preprint carries.
The Other Side: When Speed Genuinely Helped
In January 2020, early COVID-19 epidemiological preprints posted via medRxiv provided some of the first real R0 (reproduction number) estimates for the Wuhan outbreak — roughly in the 2–3 range — while very little else was known about the virus. Those early estimates were broadly consistent with later, more rigorous peer-reviewed work, and gave public health bodies real, usable numbers during the most information-starved phase of the pandemic. This is the structural benefit preprints exist to deliver — directly contrasted against Santa Clara's own structural cost of that same speed.
Registered Reports & the Center for Open Science
The Center for Open Science (COS) was founded in January 2013 by psychologist Brian Nosek and then-PhD student Jeffrey Spies, with a mission to "increase the openness, integrity, and reproducibility of scientific research" — a real, direct response to problems like Chapter 6's own reproducibility crisis.
COS's Registered Reports format changes the order of peer review itself: a journal reviews and provisionally accepts a study's introduction and methodology before data is collected or results are known. Publication no longer depends on getting a "positive," statistically significant result — which removes the structural incentive to p-hack a null finding into something publishable, or file-drawer it away entirely (Chapter 6).
Comparing every published Registered Report in psychology at the time (71 studies) against a random sample of standard psychology papers (152 studies): 96% of results in the standard literature supported the authors' original hypothesis, versus only 44% under the Registered Reports model — a 52-percentage-point drop.
Hands-On Exercises
Explain, in your own words, why the Santa Clara study's real outcome (revised and eventually published, not retracted) is a more nuanced lesson than "the preprint was simply wrong" — what does the 2021 revised figure actually tell you about the original claim?
📄 View solutionExplain, in your own words, the real structural tradeoff preprints represent, using the Santa Clara case and the January 2020 R0-estimate case as your two contrasting examples.
📄 View solutionExplain, in your own words, precisely why decoupling publication from a "positive" result is what caused the real drop from 96% to 44% in the Scheel et al. study — what does the 96% figure suggest was happening under the traditional model?
📄 View solutionChapter 8 Quick Reference
- arXiv (1991, Paul Ginsparg) pioneered preprints; bioRxiv (2013) and medRxiv (2019) extended the model, with medRxiv screening more heavily given real clinical stakes
- The April 2020 Santa Clara antibody preprint claimed a 50-85x infection undercount; critics flagged the test's own specificity confidence interval as capable of explaining nearly all "positive" results as false positives
- The study wasn't retracted — it was revised and published in 2021 at a lower, still-elevated estimate, showing peer review working after the fact rather than before public exposure
- Early January 2020 R0 preprints show the real structural benefit preprints are designed to deliver: timely, usable numbers when nothing else exists yet
- The Center for Open Science's Registered Reports (reviewed before results exist) measurably cut psychology's own suspiciously high 96% "positive result" rate down to 44%