Grounded biological data for AI drug discovery
TRANSFAC Knowledge Graph integrates 38+ years of manually curated transcriptional regulation, causal signalling and disease biology — licensed as one database to ground and validate your models.
Three curated layers, one licensed database
Delivered as MySQL tables, integrating three curated layers built and maintained for nearly four decades.
Public data is built for exploration, not for grounding production AI
01
Big data alone is not enough
LLM-style training works because text is abundant and homogeneous. Biology is neither, and drug discovery cannot absorb a wrong guess.
02
Public knowledge bases help, but fragment
Public motif and pathway databases are real curated knowledge, but fragmented, simplified and inconsistent across sources.
03
The cost is measured in years
In drug discovery the price of the wrong or incomplete grammar is measured in years and in capital, not in a bad test run.
Big Knowledge.
Robust AI for drug discovery doesn’t come from more data. It comes from comprehensive, manually curated mechanistic knowledge.
Six figures that frame what’s inside
50,950
Transcription-factor entries
118,987
Experimentally verified binding sites
1.27M
Causal reactions in TRANSPATH®
2,796
Diseases in HumanPSD™
443,444
Biomarker annotations
500+
Disease-specific reconstructed pathways
Public databases are not bad. They are insufficient for grounded AI.
Public databases
TRANSFAC Knowledge Graph
Licensed as tables, cross-referenced to the resources you already use
MySQL tables, integrating all three curated layers
Ensembl, UniProt, PubMed, Reactome, Human Protein Atlas
Maintained and versioned on a fixed release cycle
Ground, train and validate your own models on curated biology
Ensembl UniProt PubMed Reactome Human Protein Atlas
Ground your models on curated biology
License the TRANSFAC Knowledge Graph, or request the schema to see how the three layers connect.
