By MyGeneLog Team · Updated September 8, 2026 · 16 views · For classrooms
What this is. A ready-to-run statistics lesson built on real genetics, free to use and adapt in any school, college or university. Nothing to buy from us, no account, no data collected from anyone.
By 2002, genetics had a problem it did not want to talk about.
Researchers had spent a decade picking genes that looked like plausible suspects for a disease, testing whether a variant in that gene was commoner in patients than in healthy people, and publishing when the answer came out below the usual bar of p < 0.05. More than 600 positive associations had been reported this way.
Then four researchers sat down and checked which of them had held up. Of the 166 associations that had been tested three or more times by independent groups, the number that had been consistently replicated was six.
Six. Out of a hundred and sixty-six.
Nobody had cheated. Every one of those studies had used the same threshold that introductory statistics teaches, applied correctly. The threshold was the problem, and the next hour is about why.
Nothing about anybody's body is collected in this lesson, and nothing should be. This activity uses coins on purpose. Do not ask students for their own genetic results, their health, or their family's — not as an extension, not as homework, not anonymously. Count a room if you need a number; never write a name list. If a student volunteers something about their own genetics, that is their business and not the class's data.
The bar for one study is not the bar for a thousand studies, and it never was. What p < 0.05 actually promises is: if nothing is happening, a result this strong turns up about one time in twenty. Run one test and that is a reasonable risk. Run a million and it is a guarantee of fifty thousand wrong answers.
This is the whole of the multiple testing problem, and the class has just generated it from first principles rather than being told about it.
So genetics needed a number. In 2008 two groups went and found one, in two completely different ways, and this is the part students almost never get told.
One group counted. Because nearby positions in the genome are inherited together, three million tested positions are nowhere near three million independent tests. Working from the International HapMap data, they estimated the real testing burden at about one million independent tests in Europeans — and about twice that in Africans, whose genomes carry more variation and less of it travelling in blocks. Divide 0.05 by a million and you get exactly 5 × 10⁻⁸.
The other group measured. They took real genotypes from the Wellcome Trust Case-Control Consortium, shuffled the disease labels at random so that any signal found had to be noise, recorded how strong the strongest noise got, and repeated that at increasing marker densities until the answer stopped moving. Their extrapolated threshold was 7.2 × 10⁻⁸.
Two methods, one theoretical and one empirical, landing within a factor of 1.5 of each other. The field rounded to 5 × 10⁻⁸ and has used it ever since.
Worth saying plainly to a class: this is what a scientific convention looks like from the inside. Not handed down, not proved — argued out, twice, and then agreed on because agreeing was better than everybody choosing their own.
This site collects genetic associations automatically from the public GWAS Catalog and refuses anything weaker than 5 × 10⁻⁸. Here is what that filter has done, counted from its own logs.
Of 3,510 association records examined, 1,669 — 47.5% — were refused on the p-value alone. Another 1,559 were refused for other reasons: the allele data was not clean, no gene was mapped, the record was already here. 282 became pages.
Nearly half. Every one of those refused records is a published, peer-reviewed result sitting in a public database, and this site will not print it, because the threshold the class just derived says it might be a coin coming up heads five times.
The 653 variants this site does carry come from 182 different papers, and every single one of them cleared 5 × 10⁻⁸.
Give the class the figure above again and let them fight about it, because researchers are still doing exactly that.
The threshold is one number applied to everybody, and the burden is not the same for everybody. If a European-ancestry study really is running a million independent tests and an African-ancestry study is running two million, then 5 × 10⁻⁸ is the right bar for one of them and too lenient for the other. Using the stricter 2.5 × 10⁻⁸ for African-ancestry studies would be more correct and would also make it harder to publish findings from the populations that genetics has already studied least. There is no answer to that which is purely statistical.
And the number was calculated for a technology that is being replaced. A million independent tests describes genotyping chips reading common variants. Whole-genome sequencing reads rare ones too, in far greater numbers, and some argue the threshold should move to 5 × 10⁻⁹ or lower. Others answer that a threshold strict enough to be safe is strict enough to make real effects unfindable at any sample size anyone can afford.
Both arguments are live. A class with the figure and twenty minutes has the same information the people arguing have.
1. The expected count is (studies ÷ 16). The observed count will not match it, and the gap is worth a minute on its own: this is sampling variation on top of a false-discovery rate, and both are in the room at once.
2. A weighted coin would mean there was something real to find, and the discoveries would no longer all be false. The lesson depends on knowing the null is true — which is exactly what a real study never knows, and that is the point to leave hanging.
3. “How many things did they test?” Everything else follows.
4. Looking for a strong answer, not a technical one: “something that happens one time in twenty will definitely happen if you try a million times.”
5. Because the correction moved the bar below the best result the design can physically produce. This separates two ideas students routinely merge: strength of evidence is bounded by the study, and luck cannot exceed that bound.
6. 2/2ⁿ ≤ 5 × 10⁻⁸ requires 2ⁿ ≥ 4 × 10⁷, so n = 26 (2²⁶ = 67,108,864, giving 2.98 × 10⁻⁸). More flips means more participants: sample size is the only lever that moves the strongest available p-value.
7. Calculation generalises and is cheap; measurement captures the messiness of real data including the correlation structure. Agreement between two methods with different assumptions is worth more than either alone — the same logic as the vitiligo and blood pressure pages on this site.
8. Too strict: correlated tests mean fewer effective tests than markers, so dividing by the marker count over-corrects. Hence estimating independent tests. Strong students may notice this makes the threshold a property of the population, not of the technology — which is question 10.
9. No single right answer. Push for who is actually harmed: a false discovery wastes research money and can reach patients; a missed one may delay a treatment by a decade. For rare disease, missing is often the worse error, and some fields adjust for exactly that reason.
10. This is the hardest and the best. The statistics has a clean answer and the consequence does not. Do not resolve it for them. If nobody raises it, point out that the populations with the higher testing burden are the ones with the fewest studies to begin with.
11. Some real findings would certainly have been lost — several of the 6 that replicated would have failed a genome-wide bar in a study of that size. The honest answer is that the field traded discoveries for trustworthiness and decided the trade was worth it.
Use this freely. Print it, copy it, cut it up, put it on your own worksheet, translate it, change the questions. No permission needed and nothing to pay. If you credit it, mygenelog.com is enough.
If you teach with it and something in it does not work, tell us — that is worth more to us than a thank-you.
One coin per student, or a phone. A board, and a calculator for the last part. Nothing to order and nothing to spend — which is deliberate: a lesson that needs a purchase does not get taught.
No. The activity is coins and arithmetic, and everything genetic in it is explained in the text. A maths or statistics teacher can run it without having taught biology, and a biology teacher can run it without having taught probability.
None, and none should be. The activity uses coins on purpose. Do not ask students about their own genetics, health or family — not as an extension, not anonymously. Count a room if you need a number; never write a name list.
It is 0.05 divided by a million, and a million is the estimated number of independent tests in a European-ancestry genome. A separate group measured the threshold by permutation on real data and got 7.2 × 10⁻⁸. Two very different methods landing that close is why the convention stuck.