Years of schooling completed, studied genome-wide in nearly 300,000 people and then 1.1 million. A real, replicated signal exists — and even at its largest scale it explains barely a tenth of the variance in the one thing it was built to predict, in a trait shaped enormously by the society someone grows up in.
What this condition connects to
Solid lines are connections this site curates. Dashed lines mean the two ends share a research paper — worth knowing, and not a claim that one explains the other.
Prevalence
Not a categorical condition with a prevalence figure. Studied as a continuous trait (years of schooling completed) in a discovery sample of 293,723 people plus 111,349 in replication (Okbay et al. 2016, PMID:27225129), and later in 1.1 million people (Lee et al. 2018, PMID:30038396).
Inheritance
Highly polygenic: 70 common variants on this site alone from one study, among 74 total it found and 1,271 found by a later, much larger study of the same trait. Estimated to account for at least 20% of trait variation by twin/family-based estimates, while a polygenic score built from the largest GWAS to date explains only 11-13% of that variation directly — a substantial gap between total heritability and what current genome-wide data can predict.
Educational attainment — years of schooling completed — has been one of the most heavily studied traits in behavioural genetics, not because it is a clean readout of anything innate, but because it is easy to measure precisely in enormous samples. The researchers behind the largest such study say so directly: it is useful "as a proxy phenotype," a stand-in measure that correlates with harder-to-measure things, not a direct window into intelligence or ability.
Two studies, a decade apart, an order of magnitude apart
Okbay et al. 2016 combined a discovery sample of 293,723 people with a replication sample of 111,349 from UK Biobank, finding 74 genome-wide-significant loci for years of schooling completed. The associated positions were disproportionately found in genomic regions that regulate gene expression in the fetal brain — a real, biologically coherent signal, not noise. The paper's own framing of its size is the one to keep: "genetic factors are estimated to account for at least 20% of the variation," stated in the same sentence as the acknowledgment that the trait is strongly shaped by social and environmental factors. 70 of this study's loci are the variants on this page.
Lee et al. 2018 scaled the same question up by nearly a factor of four: 1.1 million people, 1,271 independent genome-wide-significant positions. That is an enormous, thoroughly replicated finding by the standards of human genetics. And the practical yield from all of it: a polygenic score built from the analysis explains 11 to 13% of the variance in educational attainment, and 7 to 10% in a related measure of cognitive performance. More than a thousand genome positions, in over a million people, and the resulting score still leaves the large majority of the variation unexplained.
Clinical detail
What this page is, and is deliberately not
This is not a test of intelligence, and no genotype on this page predicts how far any individual person will go in school. Educational attainment is diagnosed by nothing — it is not a medical condition — and it is not something a test, a guideline, or an intervention is built around here.
Three things are worth saying plainly, because this trait is more easily misread than anything else on this site. First, "years of schooling" is a measurement of an outcome, not of ability — it is shaped by family income, school access, regional and national education policy, and culture, none of which a genome-wide association study can separate out from biology. Second, even the largest version of this research — 1.1 million people, 1,271 genome positions — produces a polygenic score explaining only about a tenth of the variance in the trait it was built to predict. Third, the researchers who ran these studies use educational attainment as a "proxy phenotype" specifically because it is easy to measure at this scale, not because it is a clean stand-in for intelligence or any other single underlying thing.
None of that makes the finding fake. It is a real, replicated, biologically coherent signal — genes near the associated positions really do concentrate in fetal brain regulatory regions. What it is not, at any scale reached so far, is a usable prediction about one person.
Related variants MyGeneLog™ checks for
What a 23andMe/AncestryDNA export or raw VCF can and can't tell you about Educational Attainment comes down to these specific, well-studied positions — not a diagnosis. 77 positions are linked to this page; the ones this page's own text discusses are shown first.
No. They are associated with years of schooling completed, which the researchers who study it explicitly describe as a proxy measure, not a direct readout of intelligence. It correlates with cognitive performance but is not the same thing as it.
Can a polygenic score from these variants predict how far a specific child will go in school?
Not usefully. Even the largest study to date (1.1 million people) produces a score explaining only 11-13% of the variance in the trait — the large majority of the difference between any two people remains unexplained by it.
If genetics only explains a small share, why does this page exist?
Because the finding is real and worth stating honestly, including its own limits: a genuine, replicated, biologically coherent signal exists, and it is nowhere near strong enough to use on an individual — which is exactly the distinction a marketed "learning ability" genetic test would blur.
Does this research separate genetics from social advantage?
Not fully, and the original researchers say so themselves: educational attainment is strongly influenced by social and environmental factors. A genome-wide association study finds correlations with genotype; it does not, and cannot by itself, disentangle a family's income or a country's school system from a child's own genetic contribution.
Free to reuse. This page's text is original writing from freely-available research, licensed CC BY 4.0 — reuse it, including commercially, with attribution to MyGeneLog™. It's general research-derived information, not medical advice or a diagnosis — see Terms of Use.