By MyGeneLog Team · Updated September 8, 2026 · 11 views · For classrooms
What this is. A ready-to-run lesson on reading data critically, built on real genetics and on this site's own audit of itself. Free to use and adapt in any school, college or university. Nothing to buy from us, no account, no data collected from anyone.
There is a blood test almost everyone has had: a count of neutrophils, the white cells that deal with bacterial infection. If your count comes back below the reference range, that is neutropenia, and it is taken seriously. It can mean a drug needs stopping, a job in healthcare needs pausing, a chemotherapy dose needs reducing, or that a bone marrow biopsy is warranted.
Now the problem. Around two-thirds of Black Americans, and the great majority of people from West, Central and East Africa, carry a variant that lowers the neutrophil count in the blood without making anybody ill. It is called Duffy-null, it is one of the commonest variants in the world, and it became common because it protects against a form of malaria.
For decades the reference range those counts were compared against was built largely on people who do not carry it. Healthy people were told they had a low white cell count. Some had treatment changed or withheld over it.
Nobody made an error of reasoning. The range was applied exactly as intended. It just did not travel — and the next hour is about how to check whether a number travels before it is used on someone.
Every number below is real, from published allele frequencies. Nobody in the room is tested, counted or asked anything about themselves.
Do not turn the class into the dataset. This lesson is about populations, and the obvious wrong move is to sort the room by background and count. Do not ask students where their family is from, do not group them by it, and do not use the class as an example of any population. A room is not a sample, and a student's ancestry is not a teaching aid.
As with every lesson here: collect no genetic or health data, from anyone, by any route. The numbers in this lesson are published population frequencies and they are all you need.
The middle variant is the useful one to sit with. A prediction made in Europe would be 36 points too low for an East Asian population and 52 points too high for an African one — from the same correct arithmetic, applied to the wrong people.
And this is not an abstract variant. It sits in the gene for warfarin's target, and warfarin is a drug where the gap between too little and too much is small enough that people are tested repeatedly while their dose is found.
It would be easy to teach this as something other people got wrong. So we ran the same question at our own catalogue.
Every variant on this site names the paper it came from. We took all 187 of those papers, asked the GWAS Catalog which ancestries each study's discovery sample actually contained, and counted.
72.2% of the variants published on this site were discovered in a sample that was entirely European. Another 18.1% came from mixed samples that included Europeans. Fewer than one in ten came from a study with no European ancestry in it at all.
Cut the other way, it is starker. South Asian ancestry appears in the discovery sample of 8 of our 608 variants — 1.3%. South Asia is about a quarter of the people alive.
We are not the cause of this and we cannot fix it by ourselves: we publish what has been studied, and this is what has been studied. But a site that explains other people's blind spots while hiding its own is not teaching anything, so the number goes on the page. If a class wants to check it, every source is listed on every variant page and the GWAS Catalog is public.
Everyone agrees the imbalance is a problem. What to do about it is genuinely contested, and a class can hold the whole argument.
“Recruit more widely.” The obvious answer, and the right one in the long run. It is also the one with a history: research has been done on communities without their understanding or benefit often enough that suspicion is earned rather than irrational. Any serious answer has to include what participants are told, what they get, and who controls the samples afterwards — not only how many people are enrolled.
“Use better methods instead.” Some argue that statistical methods can borrow strength across populations and partly close the gap without new recruitment. Others answer that a method cannot invent information that was never collected, and that this argument is mostly a way of not doing the expensive thing.
“Then do not use the findings where they do not apply.” Logical, and it has a sting: applied strictly, it would mean withholding genetic risk tools from exactly the populations that have been under-studied, which is a second penalty on top of the first.
There is no consensus to hand a class. Give them the three positions and let them find out that each one has a cost.
1. It is the control. Without a case where transfer works, the lesson overshoots into “genetics never generalises”, which is wrong and unhelpful. The method has to be able to return the boring answer.
2. If the healthy people it was built from differ systematically from the patient in something that affects the measurement. Duffy-null is the cleanest example; students may also reach for age, altitude or athletic training, all of which are good.
3. The worst case is 52 points, African versus European at rs9923231. Most classes guess ten or less. Put the two numbers side by side and leave them there.
4. Carrier share is right when one copy is enough to matter; q² is right when two are required. Deciding needs the biology — is the effect dominant or recessive — which no amount of allele frequency can tell you. This is the transferable point of the whole lesson.
5. Selection acts where the pressure is. A variant that protects against a parasite becomes common where the parasite is and stays rare where it is not. Good students will connect this to lactase persistence and to why frequency maps often look like maps of something else entirely.
6. Accuracy: predictions and risk scores are less reliable outside the discovery population. Not accuracy: whose questions get asked, which diseases get studied, who benefits from the resulting drugs, and who is expected to trust a field that has not studied them. Accept any well-argued second consequence.
7. Who the participants were. Everything else — sample size, p-value, effect size — is secondary to whether the number can travel.
8. Useful but insufficient, and the honest answer is that the label is a proxy for the thing that matters, which is genetic ancestry as measured rather than as reported. Strong students may propose using measured ancestry directly, which is what modern studies do, and then run into the fact that this requires genotyping everyone.
9. The hardest practical question here. Watch for rules that quietly require racial classification of patients, and push on them: the Duffy-null case is now often handled by testing the variant itself rather than by asking about background, which is a much better rule and one a class can arrive at.
10. No right answer. The costs are unevenly distributed in all three cases, and noticing that is the outcome.
11. Strongest case against: publishing the weakness invites the finding to be quoted out of context, and could reduce trust in data that is still the best available. The answer we would give is that the number is true, discoverable by anyone with the same public API, and that a reader who finds it themselves after we hid it has learned something much worse about us. Let the class weigh it.
Use this freely. Print it, copy it, cut it up, put it on your own worksheet, translate it, change the questions. No permission needed and nothing to pay. If you credit it, mygenelog.com is enough.
If you teach with it and something in it does not work, tell us — that is worth more to us than a thank-you.
A calculator or a phone and something to write on. Every frequency is printed in the lesson, so it runs with no internet in the room.
It is a discussion about sampling, and it should stay one. The rule in the body is explicit: do not sort the room by background, do not ask students where their family is from, and never use the class as an example of a population. A room is not a sample.
Because without it a class concludes that findings never transfer, which is false. rs2075650 varies by six percentage points across five populations — a prediction made anywhere would be roughly right everywhere. The method has to be able to return the boring answer.
Yes, against the GWAS Catalog ancestry records for all 187 papers behind our variants: 72.2% of them were discovered in entirely European samples, and South Asian ancestry appears in 8 of 608. Anyone can repeat it — the sources are on every variant page and the API is public.