Open data

Gene Knowledge Graph

Most genetics references tell you what a variant is. This one records what it is joined to, and labels every join with the evidence behind it. 4,322 nodes, 4,503 edges, free to read and free to reuse.

Counted from the live catalogue API

The nodes

2,430

Variants

A position in the genome, with a cited source

1,739

Genes

Distinct gene symbols across the catalogue

83

Conditions

In plain language and in clinical detail

27

Topics

The words people search

33

Drugs

Named in a prescribing guideline

8

Supplements

With the evidence level on each

2

Sense groups

Curated smell and taste groups

One slice of it

The same picture every variant, condition and drug page on this site draws — here with real edges, not an illustration. rs713598 in TAS2R38 is shown because it reaches more different kinds of thing than any other position we hold; it is picked from the live graph each time this page is built, not chosen once and left.

rs713598 Condition: Bitter Taste Perception (TAS2R38) Bitter Taste Perception (TAS2R38) Condition Condition: Chronic Rhinosinusitis and the Bitter Taste Receptor Chronic Rhinosinusitis and the Bitt… Condition Drug: Bitter medicines Bitter medicines Drug Smell & taste: Taste Taste Smell & taste Topic: Alcohol and the flush Alcohol and the flush Topic rs713598 rs713598 TAS2R38

Solid lines are connections this site curates. Dashed lines mean the two ends share a research paper — worth knowing, and not a claim that one explains the other.

Solid lines are connections this site curates. Dashed lines mean the two ends share a research paper — worth knowing, and not a claim that one explains the other.

Multiply that by 2,430 variants and you have the graph.

The edges, and how each one is counted

An edge is one distinct (source, kind, target) triple, stated at the finest granularity its evidence actually supports. That last clause is the whole discipline: a condition page names an rsid, so that edge is variant-level; a prescribing guideline is written about a gene and never about a position, so the drug edge stays gene-level even though we know which variants sit in the gene.

Edge Count Evidence What it means
Variant → Gene 2,449 structural The variant is located in this gene, as recorded on the variant page.
Variant → Condition 1,307 curated A condition page names this variant and cites a source for it.
Gene → Drug 37 curated A prescribing guideline connects this gene to this drug. Gene-level, because that is the level the guideline is written at.
Variant → Supplement 10 curated A supplement is sold or studied on the back of this variant, with the evidence level stated.
Variant → Sense group 7 curated The variant belongs to a curated smell or taste group.
Variant → Topic 693 co-mention A paper in Europe PMC names both the topic and the variant. It does not mean one explains the other.
Total 4,503 2,449 structural, 1,361 curated, 693 co-mention

Co-mention is kept separate, deliberately. It means a paper in Europe PMC names both the topic and the variant. Researchers wrote about the two together. It does not mean one explains the other, and merging it into the curated count — which is what most graphs quietly do — is the difference between a number you can build on and a number that flatters us.

Anything that follows from two edges is reported as derived and never added to the total. 1,114 gene–condition pairs follow from the variant–gene and variant–condition edges, for example, and 29 of the 33 drug nodes have a written page so far.

What the graph leans on hardest

Degree — how many other things a node is joined to. This is a fact about this catalogue, not a claim about biology: it says what we have connected the most, which is partly a statement about what has been written up and what has not.

Genes

  1. FTO · 55
  2. MTHFR · 41
  3. TAS2R38 · 21
  4. BDNF · 20
  5. HFE · 20
  6. HLA-DQA2 · 19
  7. PITX2 · 17
  8. ADH1B · 16

Variants

  1. rs1801133 · 24
  2. rs9939609 · 21
  3. rs6265 · 19
  4. rs1801131 · 17
  5. rs1229984 · 16
  6. rs671 · 16
  7. rs1799971 · 14
  8. rs762551 · 14

What is free, and what is not

The line is not between what is interesting and what is dull. It is between a fact and its assembly.

So we are not charging for facts. We are charging for the assembly, which is the part that took the work. It is the same reason a dated snapshot costs money while the live data does not: freezing, hashing and hosting something at a permanent address is a service, and a service is a fair thing to sell.

Querying it

Grounding an AI system on it

An explicitly permitted use. The structured fields are CC BY 4.0: reuse them, including commercially, with attribution to MyGeneLog.

Three properties make it usable for that rather than merely available. Every edge names its evidence tier, so a system can weight a curated link differently from a co-mention. Every page carries the date it was last checked against its sources. And every condition records the questions we could not resolve — which is the part a system grounding on someone else's data almost never gets.

Questions about the Gene Knowledge Graph

What is the MyGeneLog Gene Knowledge Graph?

It is the whole catalogue read as connections rather than as pages: variants, the genes they sit in, the conditions they are linked to, the drugs those genes affect, the supplements and the topics — joined into one graph. As of September 2026 it holds about 4,000 nodes and about 2,000 edges. Roughly two thirds of the edges are curated, meaning a page on this site names that variant and cites a paper for it; the rest are co-mentions computed from the literature, which is a weaker claim and is labelled as one. Every edge carries which kind it is, so nothing has to be guessed from the wording. Counted and defined in one place: the Gene Knowledge Graph.

How do I query the Gene Knowledge Graph?

Through the API at /api/v1/links, which returns the edges as JSON with the variant, the gene, what it connects to, the URL of that page, and the evidence behind the link. The variant and condition endpoints are free and need no key; the full edge list is part of the Pro plan, because serving it is the expensive part. See /developers for the endpoints and the rate limits. Counted and defined in one place: the Gene Knowledge Graph.

Can I use the Gene Knowledge Graph to ground an AI system?

Yes, and it is built for it. Everything is CC BY 4.0, which permits reuse including commercially and for training or grounding, with attribution to MyGeneLog. Each edge names its evidence type, each page carries the date it was last fact-checked, and each condition records the questions we could not resolve — so a system grounding on this can tell a curated link from a co-mention and a settled figure from an open one. See /llms.txt for the machine-readable map of the site. Counted and defined in one place: the Gene Knowledge Graph.

Citing it

The counts move as the catalogue grows, so a figure taken from this page needs the date it was taken — or better, a version that cannot move. Both are available:

The structured fields are CC BY 4.0, so attribution is the whole of the licence: name MyGeneLog and link back.

What it is not

Not a diagnostic, not complete, and not a claim that these connections are causal. A curated edge means a source reported an association and a person wrote it down with the citation. Most of the variants here shift odds by a few percent and mean nothing on their own — which the pages say, one at a time, because a graph that only showed connections would leave that out.

Counts recomputed hourly from the live catalogue. Nothing on this page is typed in by hand.