Genetics papers went from one institution to nine — a survey of the 671 papers behind MyGeneLog

Genetics papers went from one institution to nine

By MyGeneLog Team · Updated September 8, 2026 · 15 views · New research

This site rests on 671 distinct papers. Every published variant and every condition page cites at least one, and those citations are the whole basis on which anything here claims to be true.

So there was an obvious question to ask of them: which universities are doing this work, and what is each of them working on?

The answer we got back was not the one we went looking for.

What we did

Papers carry author affiliations, and Europe PMC returns them. 658 of our 671 sources have at least one, 13,197 affiliation strings in all, and after matching, 599 papers yield at least one identifiable institution.

Two rules did most of the work, and both are worth stating because they decide what the numbers mean:

The first result was that everybody looked identical

Oxford, Harvard, Karolinska, Groningen, Helsinki, Copenhagen — the trait lists came back nearly the same. Atrial fibrillation. Age at menarche. Colorectal cancer. Daytime napping.

That is not what six different institutions research. It is one artefact repeated six times, and finding out why took the survey somewhere better than it was going.

The reason

A consortium paper contributing 43 variants hands the same 43 traits to every institution on its author list. If that list has sixty institutions on it, sixty institutions now appear to study exactly the same thing.

So we measured how common that is. Median number of institutions per paper, by decade of publication:

DecadePapersMedian institutionsMost on one paper
2000s132136
2010s4841171
2020s479378

The median genetics paper in our sources was a single-institution paper for two decades. In the 2020s it is a nine-institution paper. The widest one has 378 institutions on it — a saturated map of the common variants associated with human height — and it contributes zero variants to this site, because it maps a trait rather than reporting positions we index.

How much of the catalogue this affects

Paper widthPapersVariantsShare
One institution34375539.6%
2–517943522.8%
6–208027514.4%
21 or more6944123.1%

Sixty-nine papers — one in ten — carry nearly a quarter of everything the catalogue holds. And each of those sixty-nine spreads its subject matter across every institution that signed it.

A ranking here would be a ranking of consortium membership

This is the part that would have been genuinely misleading if we had published the first table and stopped.

Below is each institution's paper count, and beside it the share of those papers that are narrow — five institutions or fewer, where authorship still means a research group rather than a signature on a meta-analysis.

InstitutionPapersNarrow
Harvard Medical School5813%
Massachusetts General Hospital5330%
University of Oxford4415%
Broad Institute387%
Karolinska Institutet296%
University of Copenhagen296%
National Institutes of Health2839%

Ninety-three percent of the Broad Institute's appearances in our sources are on wide papers. For Karolinska and Copenhagen it is ninety-four. A league table built on this data would mostly be measuring which institutions join the most consortia — which is a real and interesting thing about modern science, and it is not what anybody would think they were reading.

What you can still say

Restrict to the narrow papers and the profiles separate immediately:

Those read like places with a subject. They are also, in every case, a much smaller number of papers than the headline count — deCODE has 15 narrow papers, the Institute of Cancer Research 11.

What this does not mean

It does not mean an institution absent from this list is absent from genetics.

The concrete example: the University of Texas at Austin appears in zero of our 671 sources. Texas appears plenty — UTHealth Houston on 12 papers, Baylor College of Medicine on 10, MD Anderson on 9 — but all of it is medical schools and cancer centres.

That is a fact about us. Our collector reads the GWAS Catalog, which indexes human disease-and-trait association studies. Evolutionary genetics, population genetics and computational biology do not produce the kind of paper we index, so a department strong in those is invisible here no matter how good it is. The Texas pattern — every hit a hospital — is that bias showing its shape.

So this survey measures appearance in one catalogue's source list. It is not a measure of research quality, output or importance, and any reading of it as a ranking of universities is a misreading.

One more thing, because it nearly went unnoticed

The first version of this analysis matched institution names with a pattern that required a word boundary immediately after universit. There is never a word boundary there — "University" carries on — so the pattern could not match the word it was written for. It silently dropped 262 of 671 papers, including every paper affiliated only to the University of Cambridge.

Nothing failed. No error was raised. The output looked entirely reasonable, and the top of the table was a list of hospitals. It was only obviously wrong because Cambridge appeared 288 times in the raw text and nowhere in the results.

We mention it because the same shape of error is what a reader should assume is possible in any counting exercise, including this one, and because a survey that reports its own near-miss is easier to trust than one that does not.

Why any of this matters when you read a paper

If you look up the study behind a variant on this site and find it has 171 authors from 100 institutions, that is not padding. It is what statistical power costs now. Finding a variant that shifts risk by three percent requires more people than any one university can recruit, so the recruiting is pooled and the authorship follows.

The trade is that a paper stops being the product of a place. Which is why we can tell you what the 671 papers behind this site found, and cannot honestly tell you what any one university is working on.

Every figure here was measured against this site's published sources on 8 September 2026 and moves as the catalogue grows. See where our data comes from for what those sources are.

Frequently asked questions

Does this rank universities?

No, and it should not be read as one. It counts appearances in one catalogue’s source list. Our collector reads the GWAS Catalog, which indexes human disease-and-trait association studies, so a department doing evolutionary or population genetics is invisible here regardless of how strong it is.

Why count an institution once per paper instead of once per author?

Because a 107-author paper would otherwise make one department look like forty. Counting per author measures how many people an institution sent, not whether it was involved.

Why not merge teaching hospitals into their universities?

Because whether Massachusetts General Hospital is "Harvard" is a judgement about institutional structure, and the answer differs by country. Merging them would bake that judgement into the numbers invisibly, so the rows stay separate and the reader can merge them if they want to.

Can I reproduce these numbers?

The inputs are public: the papers are the ones cited on our variant and condition pages, and the affiliations come from Europe PMC. The figures were measured on 8 September 2026 and will drift as the catalogue grows, so a recomputation later should not be expected to match exactly.

Lesson: when genetic findings don't apply to other populations
← Previous
Lesson: when genetic findings don't apply to other populations
Lesson: how to read a pedigree, and where the textbook rule breaks
Next →
Lesson: how to read a pedigree, and where the textbook rule breaks