The textbook polygenic trait, with the number usually left out: 12,111 positions, none of which means anything alone. It is also the clearest demonstration of why a genetic score built in one population predicts badly in another.
Height is the trait every explanation of "polygenic" reaches for, and it deserves the job. It is strongly inherited, it is easy to measure, and it is influenced by an enormous number of positions each contributing a fraction of a centimetre. What most explanations leave out is the actual number.
The largest study of height so far looked at 5.4 million people of diverse ancestries and identified 12,111 independent positions associated with it. Together those account for essentially all of the heritability that common variants can explain. They are not scattered evenly: they cluster into 7,209 stretches of the genome averaging about 90,000 bases each, together covering roughly 21% of it.
Read that again with a genetic report in mind. A report listing one height variant is showing you one twelve-thousandth of the picture, and this site carries two of them. That is not a flaw in the data; it is what the trait is like. Anyone selling a height prediction from a handful of positions is selling the wrong shape of thing.
Both come from a genome-wide study of population-based Korean cohorts — 8,842 people, with the promising signals re-tested in 7,861 independent Korean samples. That design matters: a signal that survives replication in a separate sample is a different class of finding from one that does not.
The same study reported something worth keeping: for height and body mass index, most of the variants it found overlapped those already reported in European samples. The positions themselves largely transfer between populations, which is a real and slightly surprising result.
Here is the part that gets misread constantly. The 12,111 positions explain about 40% of the variation in height in populations of European ancestry — where they were mostly discovered — and only about 10-20% in populations of other ancestries.
That gap is not evidence that different populations have different height genes. The same study found effect sizes, regions and prioritised genes to be similar across ancestries. What differs is how the variants are packaged: which nearby positions travel with them, and how common each version is. A score built on that packaging in one population is reading a slightly wrong map in another.
It is the single most important caveat on any polygenic score, and it applies to every trait on this site, not just this one. Prediction quality is a property of the study population as much as of the biology.
Common variants are not the whole story. Sequencing 447,461 people and replicating in 225,515 more turned up rare variants with much larger individual effects — one aggregate in the untranslated region of FGF18 is associated with height differences of up to 6 cm, which is roughly ten times what a common variant does.
And the tidy finding: 97% of those rare associations sit near regions the common-variant studies had already flagged. The two approaches converge on the same biology from opposite ends of the frequency spectrum.
Populations have grown taller over the last century faster than any gene pool can change — the answer there is nutrition, infection load and childhood health, not DNA. Within a generation, growing up well fed and healthy is the environmental part of the trait, and it is not small.
None of this is a test, and it is not how growth problems are found. A child whose growth has fallen away from their curve is a question for a doctor, who will look at growth velocity, puberty timing, nutrition and hormone levels — and none of that is answered by a genotype.
Architecture. Height is the canonical highly polygenic quantitative trait. Common SNPs are estimated to account for 40-50% of phenotypic variance, and the 2022 GWAS meta-analysis of 5.4 million individuals resolved 12,111 independent genome-wide significant SNPs that capture nearly all of that common-variant heritability. Those SNPs fall in 7,209 non-overlapping segments with a mean span of about 90 kb, covering approximately 21% of the genome; density is uneven and the denser regions are enriched for biologically relevant genes.
Portability. Out-of-sample estimation gave 40% of phenotypic variance explained in European-ancestry populations using the 12,111 SNPs (45% using all HapMap 3 SNPs), against roughly 10-20% (14-24%) in populations of other ancestries. Effect sizes, associated regions and prioritised genes were similar across ancestries, so the loss is attributed to linkage disequilibrium structure and allele-frequency differences within associated regions rather than to different underlying biology. Any polygenic score for any trait inherits this limitation.
The two rows here. rs6918981 (HMGA1) and rs13273123 (PLAG1) come from a Nature Genetics 2009 genome-wide study of Korean population-based cohorts: 8,842 samples for discovery with replication of promising signals in 7,861 independent Korean samples across eight quantitative traits. For height and BMI specifically, most detected variants overlapped those previously reported in European samples. Per-allele effect estimates for height variants are on the order of a few millimetres and are not clinically interpretable in an individual.
Rare-variant contribution. Whole-genome sequencing in 447,461 UK Biobank participants with replication in 225,515 All of Us participants identified 90 rare and low-frequency single-variant associations across height, BMI and waist-hip ratio, plus 135 coding and 51 non-coding aggregate associations; a 5'UTR aggregate in FGF18 reaches effects of up to 6 cm on height. 97% of rare-variant associations lie near loci already identified by common-variant GWAS, and ultra-rare variants explain a small fraction of heritability relative to common variants.
Clinical boundary. Nothing here concerns monogenic short stature or overgrowth syndromes, skeletal dysplasias, growth hormone deficiency or the endocrine and nutritional causes of growth failure. Those are diagnosed from growth velocity, bone age, puberty staging and targeted testing. A common height variant has no role in that assessment.
What a 23andMe/AncestryDNA export or raw VCF can and can't tell you about Height comes down to these specific, well-studied positions — not a diagnosis.
Not usefully. The best available score uses 12,111 positions and explains about 40% of the variation in the populations where it was built — and only about 10-20% elsewhere. A test reading a handful of positions explains far less than that. Mid-parental height and a growth chart do better, and cost nothing.
Essentially nothing on their own. They are real, replicated associations, and each moves the average by a few millimetres across a population. Between them they are two positions out of more than twelve thousand.
Because the scores were mostly built in European-ancestry samples. The same regions and similar effect sizes show up across ancestries, but which nearby positions travel together, and how common each version is, differs — so a score reads a slightly wrong map. It is a limitation of the studies, not a difference in biology.
Because the environmental share is large and it changed. Nutrition, childhood illness and general health moved a long way in a century, and the gene pool did not. Strongly heritable within a population and strongly environmental between generations are not contradictory — they are different questions.
Free to reuse. This page's text is original writing from freely-available research, licensed CC BY 4.0 — reuse it, including commercially, with attribution to MyGeneLog. It's general research-derived information, not medical advice or a diagnosis — see Terms of Use.