Reproducibility: Can Other Scientists Get the Same Result?

Scientific Methodology

Chapter 6 · Reproducibility: Can Other Scientists Get the Same Result?

A single published result, however striking, isn't the end of the scientific process — it's a claim waiting to be checked. This chapter covers real, hard evidence of how often that check fails, and a real, influential theory of why.

A Real, Landmark Test

The Open Science Collaboration, Science, 2015

A large team of researchers attempted to replicate 100 real psychology studies, originally published in 2008 across three prestigious journals. In the original studies, 97% reported a statistically significant effect. In the real replications — using the original materials wherever possible — only about 36% produced a statistically significant effect, and average effect sizes were roughly half the size originally reported.

This single result became the most cited real data point in what's now widely called the reproducibility crisis — and it wasn't confined to psychology; similar patterns have since been documented across cancer biology, economics, and other fields.

Why This Happens: A Real, Influential Framework

Physician-researcher John Ioannidis's real 2005 PLOS Medicine paper, "Why Most Published Research Findings Are False," argued the problem is structural, not simply a matter of individual dishonesty. He identified real, specific factors that make a published finding less likely to be genuinely true:

Small Sample Sizes

Smaller studies produce noisier, less reliable real estimates, more prone to chance findings.

Small Effect Sizes

A genuinely small real effect is harder to distinguish reliably from statistical noise.

Publication Pressure

Novel, striking, significant results are far more publishable than a careful replication or a null result.

A Real, Structural Incentive Problem
The Open Science Collaboration's own real interpretation pointed toward exactly this cultural pattern: academic incentives reward novelty over reproduction. A researcher who spends real time and resources carefully replicating someone else's already-published finding gains far less career credit than one who publishes something new — even though replication is what actually confirms a finding is real.

The Real, Measured Drop

Original studies (2008)Real replications (2015)
Statistically significant result97%~36%
Average effect sizeBaselineRoughly half
The Real, Practical Lesson
A single striking published result — especially from a small study in a competitive, "hot" research area — deserves real, genuine caution until independently replicated. Chapter 8 covers real, current reforms (preprints, open data, pre-registration) built directly in response to this exact crisis.

Hands-On Exercises

Exercise 1

Explain, in your own words, why a drop from 97% to 36% significant results is a more alarming real finding than it might first appear — what does it suggest about how many of the original 2008 findings were likely real effects at all?

📄 View solution
Exercise 2

Explain, in your own words, how Ioannidis's real "small sample size" and "publication pressure" factors could combine to produce a specific, concrete false finding — walk through a plausible real scenario.

📄 View solution
Exercise 3

A classmate argues the reproducibility crisis proves science itself can't be trusted. Explain, in your own words, why the existence of the 2015 study itself is actually evidence against that conclusion.

📄 View solution

Chapter 6 Quick Reference

  • The real 2015 Open Science Collaboration study replicated 100 psychology studies; only ~36% reproduced a significant effect, down from 97% originally, with effect sizes roughly halved
  • Ioannidis's real 2005 paper identified small sample sizes, small effect sizes, and publication pressure as structural drivers of unreliable findings
  • Academic incentives reward novel findings far more than careful replications
  • A single striking result deserves real caution until independently replicated