For AI drug-discovery teams

Grounded biological data for AI drug discovery

TRANSFAC Knowledge Graph integrates 38+ years of manually curated transcriptional regulation, causal signalling and disease biology — licensed as one database to ground and validate your models.

Three curated layers, one licensed database

Delivered as MySQL tables, integrating three curated layers built and maintained for nearly four decades.

TRANSFAC®

Transcription factors with experimentally verified binding sites and the most comprehensive library of DNA motifs.

Layer 01

TRANSPATH®

Causal, directional signal-transduction reactions across the proteome, including modified forms and complexes.

Layer 02

HumanPSD™

Diseases, clinical trials and the biggest collection of biomarkers and drug targets with mechanistic annotations.

Layer 03

Why grounding

Public data is built for exploration, not for grounding production AI

01

Big data alone is not enough

LLM-style training works because text is abundant and homogeneous. Biology is neither, and drug discovery cannot absorb a wrong guess.

02

Public knowledge bases help, but fragment

Public motif and pathway databases are real curated knowledge, but fragmented, simplified and inconsistent across sources.

03

The cost is measured in years

In drug discovery the price of the wrong or incomplete grammar is measured in years and in capital, not in a bad test run.

The alternative has a name

Big Knowledge.

Robust AI for drug discovery doesn’t come from more data. It comes from comprehensive, manually curated mechanistic knowledge.

The numbers

Six figures that frame what’s inside

50,950

Transcription-factor entries

118,987

Experimentally verified binding sites

1.27M

Causal reactions in TRANSPATH®

2,796

Diseases in HumanPSD™

443,444

Biomarker annotations

500+

Disease-specific reconstructed pathways

Curated vs public

Public databases are not bad. They are insufficient for grounded AI.

Public databases

Fragmented motif and feature collections

Mixed genes and protein features

Static snapshots, inconsistent across sources

Correlation-level associations

TRANSFAC Knowledge Graph

One integrated, cross-referenced graph

Genes, factors, reactions and disease joined

Maintained on a fixed curation cycle

Causal mechanism, traceable to the experiment

Delivery

Licensed as tables, cross-referenced to the resources you already use

Format

MySQL tables, integrating all three curated layers

Cross-references

Ensembl, UniProt, PubMed, Reactome, Human Protein Atlas

Updates

Maintained and versioned on a fixed release cycle

Use

Ground, train and validate your own models on curated biology

Ensembl UniProt PubMed Reactome Human Protein Atlas

Explore the platform

Three ways in, one body of knowledge

TRANSFAC Workspace

Run the whole regulatory analysis yourself, on one no-code platform.

Open

TRANSFAC Services

Give us the target or disease; our scientists deliver the answer.

Open

Overview

See all three ways in and pick the one that fits your team.

Open

Ground your models on curated biology

License the TRANSFAC Knowledge Graph, or request the schema to see how the three layers connect.