By MyGeneLog™ Team · Updated September 9, 2026 · For classrooms
Next Generation Science Standards
Adopted verbatim by 20 states and DC, adapted by 25 more — roughly 45 states in all.
HS-LS3-1
Grades 9–12
LS3: Heredity — Inheritance and Variation of Traits
Ask questions to clarify relationships about the role of DNA and chromosomes in coding the instructions for characteristic traits passed from parents to offspring.
2022 revised national curriculum (Korea)
The national curriculum for every school in South Korea.
9과21-05
Grade 9 Science (Korea, middle school year 3)
Reproduction and heredity
사람의 유전 형질과 유전 연구 방법을 알고, 가계도를 분석하여 사람의 유전 현상을 설명할 수 있다.
Our translation Know human genetic traits and the methods used to study human inheritance, and explain human inheritance by analysing a pedigree.
Six states use neither the NGSS nor standards derived from it — Florida, North Carolina, Ohio, Pennsylvania, Texas, Virginia. If you teach in one, read the standard text above rather than the code.
What this is. A ready-to-run genetics lesson, free to use and adapt in any school, college or university. Nothing to buy from us, no account, no data collected from anyone.
Everyone alive owns a copy of the same book.
That is the lesson's whole vocabulary. Everything below is either a closer look at one of those five lines, or the thing students actually do with them.
The one sentence to keep. A gene is a chapter; an rsID is the number given to a single letter inside it. One chapter holds many such letters — which is why "I have the BRCA1 gene" cannot be right on its own. Everybody has BRCA1. What differs is which letters are in their copy of it.
This is the commonest confusion in the whole topic, and it is worth two minutes before the picture rather than an argument after it. The word "two" is about to mean two different things.
The figure below is the first kind: two different people, one letter apart. Keep the second kind in mind, because it is the reason a result has two letters in it.
| In the book | In a genome | What it is |
|---|---|---|
| The letters, and the paper they are printed on | DNA | Four letters — A, T, G, C — paired into a ladder and twisted. That twist is the famous double helix. It is the paper, not the writing. |
| The 23 volumes it is bound into | Chromosomes | Two metres of DNA wound tight enough to fit inside a cell you cannot see. You carry 23 pairs. |
| A chapter that explains how to make one thing | Gene | A stretch of the text carrying the instructions for one product, usually a protein. About 20,000 of them. |
| A letter that reads differently in one person’s book than in another’s | SNP | A single position where people differ. Almost every letter is identical in everyone; these are the rare ones that are not. |
| The catalogue number for that exact spot | rsID | A permanent name so that everybody, everywhere, means the same position. rs1815739 is the same place in every paper on Earth. It names the place, not the letter you happen to have there. |
So the containment runs one way: a chapter contains many such positions, and a position sits inside at most one chapter — often none at all, because most of the book is not chapters.
Where the analogy breaks, and it matters. A book is read front to back and says the same thing to every reader. This one is not. Different chapters are read in different cells — a muscle cell and a liver cell read the same book and use different parts of it. Most of a chapter's text is cut out before the instruction is used. And one chapter can yield several different products depending on which parts are kept. Push the metaphor past "chapter and letter" and it will start telling you things that are not true.
A gene is a source file; an rsID names a line inside it that differs between forks. The same caution applies and applies harder: code is executed as written, top to bottom, and DNA is not.
The numbers are worth saying out loud, because they are not the numbers most people expect.
| Thing | Roughly how many, in one human |
|---|---|
| Chromosomes | 46, in 23 pairs |
| DNA letters | about 3 billion per copy |
| Length of DNA in a single cell, unwound | about 2 metres |
| Protein-coding genes | about 20,000 |
| Genes that are transcribed but make no protein | tens of thousands more |
| rsIDs described so far, across all people | hundreds of millions |
The 20,000 is the surprise. Before the genome was sequenced, serious estimates ran to 100,000 and higher, on the reasoning that a complex organism needs a lot of instructions. The count came back far lower, and the resolution is that the instructions are reused: one gene can yield several products, and a great deal of the genome is regulation deciding when the rest runs.
This is the part where the two words stop being definitions and start being something you can measure. Everything here uses the free pages on this site; nothing is submitted, and no account is needed.
Before you start — the one rule. Nobody brings their own genetic data to this lesson, and nobody is asked whether they carry anything. There is no test here and no result about any person in the room. The activity is about a public catalogue, not about you. If a student asks about their own DNA, the honest answer is that a class is the wrong place to find out and a doctor or genetic counsellor is the right one.
Write two numbers on the board, agreed by the class before any searching:
Almost every class predicts the famous gene will have more. Write the prediction down. The prediction being wrong is the lesson, and it only works if it is committed to first.
Search each gene symbol on this site and count the variant pages returned. As of writing, the catalogue holds 2,976 variants across 2,076 genes, and the counts look like this:
| Gene | Known for | rsIDs here | What those entries are about |
|---|---|---|---|
| PITX2 | heart rhythm | 15 | atrial fibrillation — all fifteen |
| FTO | body weight | 10 | body mass and related measures |
| ABO | blood group | 9 | a range of traits |
| APOE | Alzheimer risk | 4 | cholesterol and Alzheimer disease |
| TAS2R38 | bitter taste | 3 | tasting PTC |
| MTHFR | folate processing | 2 | folate metabolism |
| CFTR | cystic fibrosis | 2 | cystic fibrosis |
| ACTN3 | muscle fibre type | 1 | muscle fibre type |
| TP53 | the most-studied gene in cancer | 1 | basal cell carcinoma |
| BRCA1 | breast and ovarian cancer risk | 1 | age at first period |
Read the last column before going on. It is doing more work than the counts are.
Put the prediction beside the result. TP53 is the most-studied gene in all of cancer biology and it has one entry here. PITX2 has fifteen. Now the question that makes this a science lesson rather than a browsing exercise: does PITX2 vary more between people than TP53 does?
It does not. The count measures how this catalogue was built. Almost everything here comes from genome-wide association studies, which scan the genome for positions that differ between people who have a trait and people who do not. Atrial fibrillation has been studied that way many times, PITX2 comes up every time, and each study leaves entries behind. All fifteen PITX2 rows are the same trait, found again and again.
The two bottom rows make the point sharper than any explanation could.
TP53 is arguably the most important gene in cancer research, and its single entry here is for basal cell carcinoma. That is not an error. TP53 matters mostly through mutations that arise in a tumour during a person's life, which are not inherited and do not show up in this kind of study at all.
BRCA1 is the gene everybody in the room has heard of, and its one entry here has nothing to do with cancer. It is about the age somebody has their first period. The cancer-causing BRCA1 changes are rare, individually powerful, and found by sequencing families — a completely different method from the one that built this catalogue. So the famous cancer gene appears in a genetics catalogue under a trait nobody associates with it, because that is the only kind of variant the method could see.
The count, and even the subject of an entry, reflects which questions have been asked and with what instrument — not how variable a gene is or what it matters for. A student who takes that one idea away has learned something that applies to every dataset they will ever meet.
Compare the sizes honestly. This site holds about 3,000 variants across about 2,000 genes. There are roughly 20,000 protein-coding genes, and hundreds of millions of rsIDs described. So this catalogue covers around a tenth of the genes and a tiny fraction of one percent of the positions.
Worse than small, it is unevenly small. The studies it is built from were run disproportionately in people of European ancestry, so a position that matters in other populations is more likely to be missing — not because it does not exist, but because fewer studies went looking. That is a limitation of the source material, and it is the reason absence of a variant from any catalogue is never evidence of absence in a person.
Given all that, how much should anyone trust a count of variants per gene?
One position: these counts are an artefact of study design and should never be presented as a property of a gene. Another: attention is not random — a gene that keeps turning up across independent studies of different traits probably is doing something broadly important, and the count carries real signal buried in the bias.
Both are argued in the literature, and the disagreement is not settled. Students can hold the same evidence the field holds, which is the point of showing them the raw counts rather than a conclusion.
1–4 check the containment relationship and separate the physical structure (DNA, helix, chromosome) from the informational unit (gene) and the label (rsID). Question 4 catches the commonest confusion in the room: students routinely say the gene is a helix. The helix is the material the gene is written on.
5 and 6 are the centre of the lesson. Accept any two mechanisms — more studies, different study type, gene length, how variable the region is — but push hard on the second half. The evidence that separates them is the number of published studies per gene, which is countable, and students should notice that this makes the question empirical rather than a matter of opinion.
6 has a specific answer worth landing: the catalogue was built from studies of common variants, and the BRCA1 changes that cause cancer are rare ones found by sequencing families. A method that only sees common differences will catalogue a famous gene under whatever common difference it happens to carry. Students who reach this on their own have understood what an instrument is.
7 is the safety question and the one most likely to come up unprompted. The answer is that a consumer array tests a small selected panel; absence from the panel is not absence in the person. Do not let this drift into interpreting anyone's actual result.
8 rewards students who can hold a model and its limits at once. Good answers: the book is not read front to back; different cells read different chapters of the same book; most of a chapter is cut out before the instruction is used; one chapter can make several products. Students who have written code often reach for the file-and-line version in the aside — accept it, and hold them to the same standard about where it fails.
9 is about who the data was collected from. Students often assume a catalogue is a neutral mirror of nature. It is a record of what was funded and studied.
10–12 are for students who want a research problem. Question 12 has a concrete real answer — coordinates shift between genome builds while rsIDs are meant to remain stable, and mismatched builds are a genuine source of error in practice.
Every gene and position named here has a page on this site, free and with its sources listed. Good places to continue:
Free to use, copy and adapt in any classroom, at any level, including commercially. Credit is welcome and not required. Nothing on this site collects anything from a student: no account, no upload, no genetic data, ever.
© 2026 MyGeneLog™. Licensed under CC BY-NC 4.0 — share or adapt this article, including translations, with credit to MyGeneLog™ and a link back to this page. Commercial use is not permitted without our written permission. Full terms: mygenelog.com/terms.
Licensed CC BY-NC 4.0 — when citing this, name MyGeneLog™ and link to this exact page. Commercial use needs our written permission.
Lesson: genes and rsIDs, and what a catalogue actually measures. MyGeneLog™. https://www.mygenelog.com/updates/lesson-genes-and-rsids-as-code
A gene is a stretch of DNA carrying the instructions for one product, usually a protein. An rsID names a single position where people differ. One gene contains many rsIDs; an rsID sits inside at most one gene, and often none, because most of the genome is not gene.
About 20,000 protein-coding genes, plus tens of thousands more that are transcribed without making a protein. The first estimates ran to 100,000 and higher; the count came back lower because the instructions are reused and much of the genome is regulation.
No. The double helix is the shape of DNA itself — the material a gene is written on. A gene is a stretch of that material. A chromosome is DNA wound tightly so two metres of it fits inside one cell.
No, and it must not use any. Nobody brings their own genetic data, nobody is asked what they carry, and no result about any person in the room is produced. The activity uses a public catalogue.