For AI drug discovery teams

Grounded biological data for AI driven research

TRANSFAC Knowledge Graph integrates 38+ years of manually curated transcriptional regulation, causal signaling, and disease biology, delivered as MySQL tables licensable for AI training.

Ns Hero Kg Dark

Curated knowledge to ground AI-driven research

It provides curated, mechanistically connected biological knowledge to ground AI models and support AI-driven research across four key areas.

Drug Discovery & Development

Discover targets, biomarkers, mechanisms, and therapeutic opportunities.

Biotechnology & Synthetic Biology

Understand and engineer gene regulation and biological systems.

Precision & Personalized Medicine

Interpret patient-specific molecular data for stratification and treatment decisions.

Multi-Omics & Disease Mechanism Research

Connect multi-omics data to causal pathways and disease mechanisms.

The schema

The basic MySQL schema

The basic MySQL schema of the TRANSFAC® Knowledge Graph: one entity joined to its regulatory layer, interactions, disease context, annotation and references.

The basic MySQL schema of the TRANSFAC® Knowledge Graph: one entity joined to its regulatory layer, interactions, disease context, annotation and references. Tap to enlarge.

The full schema is available under NDA

The full MySQL schema is available upon request following the signing of an NDA.

Request TRANSFAC Knowledge Graph schema

A simple example

Walk through a simple example to see how the TRANSFAC Knowledge Graph connects transcriptional regulation, causal signaling, and human health insights in one unified view.

TRANSFAC® for who binds where, TRANSPATH® for how molecules interact and signal, HumanPSD™ for what it means for human health and therapy.

TRANSFAC® for who binds where, TRANSPATH® for how molecules interact and signal, HumanPSD™ for what it means for human health and therapy. Tap to enlarge.

Where Big Knowledge strengthens AI-driven research

Curated knowledge added to complex research data

TRANSFAC Knowledge Graph strengthens the central stages of the AI-driven research pipeline by adding curated, mechanistically connected biological knowledge to complex research data. It helps integrate and interpret omics, literature, clinical, and other data through established knowledge of transcriptional regulation, causal signalling, disease biology, biomarkers, and drug targets. This provides AI models with more than statistical associations: it adds biological context, causal structure, and traceable evidence. The result is better-grounded model training, more meaningful biological interpretation, stronger target and biomarker discovery, and more reliable decisions across drug discovery, biotechnology, precision medicine, and multi-omics research.

The AI-driven research pipeline, with the TRANSFAC Knowledge Graph as the Big Knowledge layer feeding data integration and biological interpretation.

The AI-driven research pipeline, with the TRANSFAC Knowledge Graph as the Big Knowledge layer feeding data integration and biological interpretation. Tap to enlarge.

Why grounding

Public data sets are built for exploration and not for grounding production AI

01

Big data alone

Large language model style training works because text is abundant and homogeneous. Biology is neither. Drug discovery datasets are small relative to the dimensionality of the systems they describe, and the raw omics are noisy. Models trained on raw omics from scratch learn statistical correlations and miss causation, which is why they look strong on benchmarks and fail on a new disease subtype, cell type, or patient cohort. Omics datasets alone give a model the vocabulary of biology without the grammar. Models trained on them can speak fluently and still be wrong.

02

Public knowledge bases

Adding public motif and pathway databases helps. They are real curated knowledge. But they are fragmented, simplified, and missing the cell type and tissue-specific detail that drug discovery actually depends on. It’s the difference between learning a language from a tourist phrasebook and learning it from a proper grammar.

In drug discovery the cost of the wrong or incomplete grammar is measured in years and in capital, not in test set accuracy.

Strong AI starts with strong biological knowledge: when the biological data foundation is reliable, every stage above it becomes more trustworthy.

Strong AI starts with strong biological knowledge: when the biological data foundation is reliable, every stage above it becomes more trustworthy. Tap to enlarge.

The alternative has a name

Big Knowledge.

Robust AI for drug discovery doesn’t come just from more data. It comes from comprehensive, manually curated molecular biology — the kind that took 38 years to build.

Three curated layers. One licensed database.

A licensed biological knowledge graph delivered as MySQL tables, integrating three curated layers built and maintained by the geneXplain team over 38 years. Every entry links to its primary literature. Every reaction has a direction. Every annotation carries an evidence class. Updates ship twice a year.

Layer 01 — TRANSFAC

Transcription factors with their experimentally verified binding sites and the most comprehensive library of DNA motifs — the foundation of regulatory genomics for nearly four decades.

Layer 02 — TRANSPATH

Causal, directional signal transduction reactions across the proteome including modified forms and protein complexes. Not correlation networks — curated reactions with direction, cellular context and source.

Layer 03 — HumanPSD

Diseases, clinical trials, the biggest collection of biomarkers and drug targets with mechanistic annotations and evidence classes — connecting molecular biology to clinical context.

Cross-references resolve to Ensembl, UniProt, PubMed, Reactome, and Human Protein Atlas.

The numbers

Six numbers that frame what’s inside

50,000+

Transcription factor entries across species

100,000+

Experimentally verified binding sites

1.27M

Causal reactions in TRANSPATH

2,700

Diseases in HumanPSD

400,000+

Biomarker annotations

500+

Disease-specific reconstructed pathways

Recent database statistics across the TRANSFAC® Knowledge Graph.

Recent database statistics across the TRANSFAC® Knowledge Graph. Tap to enlarge.

Vs. public databases

Public databases are not bad, they are insufficient for grounded AI

Public databases often give you fragmented sequence motif collections, mixed genes or protein features, statistical correlations, and aggregated network edges. TRANSFAC Knowledge Graph gives you curated entries traceable to the experiments that produced them, reactions with a direction and cellular context, signaling pathways reconstructed from primary literature, and disease predictive and prognostic biomarkers and drug targets.

  • High quality manually curated database (>1000 person years curation efforts)
  • Causal, directional reactions (not correlation)
  • Primary-literature provenance on every entry
  • Mechanistic disease biomarker classification
  • Signal transduction pathways
  • Curated updates twice a year
  • Licensable for AI training

TRANSFAC Knowledge Graph advantages

Requirement by requirement, against the public alternatives.

Data requirement

TRANSFAC Knowledge Graph

Public motif databases

Public pathway / network databases

Curated TF binding motif collection

>11,000 expert curated PWMs derived from experimentally validated TFBS with manually optimized alignments ensuring high biological fidelity

High quality open motif collections with broad coverage; typically fewer profiles and less harmonization across experiments

Not applicable

Experimentally verified TF binding sites

Extensive curated TFBS dataset with links to TFs, genes, species, experimental context, and literature evidence

Often focused on motif models or ChIP-derived regions; experimental site-level annotation is more limited or heterogeneous

Not applicable

Composite regulatory elements (TFBS combinations)

Unique strength

Curated and computationally derived combinations of TF binding sites (composite elements) capturing cooperative and combinatorial regulation — critical for real gene control logic

Typically represent individual motifs; limited or no systematic modeling of TFBS combinations

Not applicable

Context-specific motif collections

Ready-to-use tissue-, cell-type-, and disease-specific motif and TFBS collections reflecting biological context

Context annotations exist but are not typically delivered as structured, ready-to-use regulatory layers

Not applicable

Genome-wide TFBS predictions

Genome-wide TF binding maps across ~1,000 TFs, >300 cell types, and >50 tissues using PWMs and MEALR models

Genome-wide predictions available but usually less context-aware and less focused on combinatorial regulation

Not applicable

Enhancers and silencers (context-specific regulatory elements)

Expert predicted genome-wide enhancers and silencers specific to tissues, cell types, diseases, and phenotypes, based on regulatory grammar and TF combinations

Enhancer datasets exist (often experimental), but are typically not integrated with TF combinatorial logic or mechanistic regulatory modeling

Not applicable

Causal signaling reactions

Large scale collection of curated causal, directional signaling reactions suitable for mechanistic modeling

Not applicable

Strong pathway resources exist, but may include mixed evidence types (causal and associative)

Mechanistic molecular detail

Explicit representation of genes, proteins, isoforms, post-translationally modified forms, and protein complexes, enabling true mechanistic resolution

Not applicable

Pathway resources provide structured reactions but often simplify molecular states or aggregate entities

Protein complexes and modified forms

Detailed modeling of complexes and molecular states within signaling and regulatory processes

Not applicable

Present in curated pathways but depth and consistency vary

Protein Genome Map (protein-centric genome annotation)

New layer

Mapping proteins, their isoforms, modifications, and functional states back onto genomic regulatory regions, bridging genome and proteome in one framework

Not available

Not available

Disease-specific pathways

Expert-reconstructed disease pathways derived from curated biomarkers and causal signaling networks

Not applicable

Disease pathways exist but are often generalized or not reconstructed via causal graph approaches

Disease biomarkers

Structured biomarker knowledge classified by causality, mechanism, prognosis, and drug relevance

Not applicable

Disease associations present but not organized as a dedicated mechanistic biomarker layer

Disease similarity maps

Disease similarity networks based on shared biomarker profiles with expert-weighted evidence types

Not applicable

Disease relationships exist but usually not based on structured biomarker similarity modeling

Integration across biological layers

Unified schema connecting DNA motifs → TFs → regulatory modules → signaling pathways → disease biology → drugs

Typically focused on regulatory layer only

Typically focused on pathway/network layer only

Literature traceability

Each entry linked to primary literature with clear evidence annotation

References provided but depth varies

Curated resources provide references; large scale networks may include predicted associations

Data consistency and curation depth

>38 years of continuous expert curation with consistent schema and harmonized biological representation

Valuable open resources but variable consistency and depth

Strong curated resources exist alongside aggregated datasets with mixed evidence types

Best use case

Mechanistic modeling, AI training, causal inference, and regulatory design

Motif discovery, benchmarking, exploratory analysis

Pathway enrichment and general network analysis

Honest limitation

Commercial licensed dataset optimized for depth, consistency, and mechanistic modeling

Open and accessible; widely used for benchmarking

Broad and accessible; may require integration and filtering for mechanistic use

Track record

Who already builds on this

A growing number of AI and biotechnology companies use geneXplain’s curated regulatory, signalling and disease knowledge to strengthen drug discovery workflows — from target and biomarker discovery to mechanism-of-action analysis and therapeutic hypothesis generation.

In production

One example is a leading AI drug discovery company that licensed the full TRANSFAC Knowledge Graph in 2025 for foundation model training and biological grounding.

The database provides its AI models with curated, mechanistically connected knowledge of gene regulation, causal signalling and disease biology.

Beyond commercial AI applications, the knowledge underlying TRANSFAC, TRANSPATH and HumanPSD has also supported peer-reviewed research in master regulator discovery, disease-mechanism reconstruction and drug repurposing.

References

1.  Myer et al., Gastro Hep Adv (2022).

2.  Kel et al., BMC Bioinformatics (2019).

3.  Lloyd et al., Disease Models & Mechanisms (2020).

4.  Kolmykov et al., Nucleic Acids Research (2020).

Where it changes your results

Where TRANSFAC Knowledge Graph changes your results

What makes the difference is not the number of motifs or pathways, but the ability to represent how they work together: as composite regulatory elements, context-specific enhancers, and mechanistically resolved molecular states.

01

Foundation model grounding

Models trained on raw omics or aggregated public data tend to learn correlations that do not generalize. Grounding your model in curated regulatory, signaling, and disease knowledge adds causal structure, biological constraints, and traceability to primary literature.

02

Mechanism-based target discovery

Expression-based approaches identify associations. Causal upstream modeling connects disease phenotypes to transcriptional master regulators through signaling pathways, enabling identification of actionable targets rather than correlated markers.

03

Multi-layer integration

Instead of stitching together separate motif, pathway, and disease resources, all layers are already connected in a single schema. Regulators, composite elements, signaling reactions, biomarkers, and diseases are linked consistently, enabling coherent mechanistic interpretation across relevant omics layers.

04

Synthetic regulatory module design

Regulatory design requires more than individual motifs. Using experimentally anchored binding sites, composite regulatory elements, and context-specific enhancer logic enables the design of promoters, enhancers, and regulatory circuits that reflect real biological control mechanisms.

The four places the Knowledge Graph changes a result: grounding, target discovery, multi-layer integration and regulatory design.

The four places the Knowledge Graph changes a result: grounding, target discovery, multi-layer integration and regulatory design. Tap to enlarge.

Delivery

Everything that ships with a license

  • MySQL database containing TRANSFAC, TRANSPATH, and HumanPSD tables, including core entity tables and the cross-reference linking tables that connect them
  • Schema documentation and data dictionary, field-level, in PDF
  • Loader scripts for standard MySQL deployment
  • Cross-references to Ensembl, UniProt, PubMed, Reactome, and Human Protein Atlas
  • Experimental evidence annotations on regulatory entries, binding sites, and reactions
  • Half a year updates with changelog and migration notes
  • Contract-defined SLA for technical support, including database schema guidance, data ingestion assistance, and update integration
  • Delivery: within 10 business days of access being granted

What you license and what you don’t

Not included

Derived position-weight matrices (PWMs)

HMM models

Combinatorial modeling based on sparse logistic regression (MEALR)

Technical tables required by the geneXplain GUI software

Included

TRANSFAC, TRANSPATH and HumanPSD MySQL tables

Cross-references to Ensembl, UniProt, PubMed, Reactome and Human Protein Atlas

Experimental evidence annotations

Some assets that customers sometimes assume are bundled with TRANSFAC are separate licensed products.

Licensing

Two license models. Scoped options on request.

License model 01 — Controlled AI License

Full database delivery for internal AI training, validation (fact checking) and research. Restrictions: no database reconstruction, no API exposure of core services.

Full MySQL database download

Train models on the full corpus, internally

Outputs of trained models are commercializable

Request pricing

License model 02 — AI & Externalization License

Full training rights plus deployment, integration, and externalized AI services. Restrictions: no replication or redistribution of the database itself.

All Controlled AI rights, plus:

Deploy trained models in commercial products and services

Integrate model outputs into customer-facing AI features

Request pricing

Disease area slice

Scoped licenses for oncology, inflammation, neurology, or rare diseases — for teams whose pipeline doesn’t need the full corpus.

Contact for scope

FAQ

Frequently asked questions

The database is manually curated from primary literature with traceable citations on every entry. Reactions are causal and directional, not correlative. Disease pathways are reconstructed by curators, not aggregated from network edges. Updates are each half a year, with a changelog. Public databases are valuable for exploration. The database is built for grounding production AI.

The databases undergo structured expert curation and are released on a biannual basis. Each release is accompanied by a detailed changelog and migration documentation, supporting traceability, reproducibility, and consistent downstream use. The curation process is maintained by a continuously active scientific team with more than 38 years of accumulated domain expertise.

The licensee owns all model outputs and trained weights produced under the license. The license covers the right to use the database, it does not claim downstream IP.

The data is curated from peer-reviewed scientific literature and fully traceable to primary sources. GeneXplain applies strict curation standards and provides the data under clearly defined commercial terms. Given the nature of scientific publishing and data aggregation, indemnification is provided within reasonable commercial limits and specified explicitly in the license agreement.

Yes. Request the TRANSFAC Knowledge Graph schema or book a discovery call to get access to an NDA-gated demo dataset.

Within 10 business days of access being granted, per the standard delivery exhibit.

Talk to the team that built it

Bring your specific question. We’ll tell you whether the database fits.

The database has been developed and curated for 38+ years by the team that originated TRANSFAC. Bring a target class, a disease area, or an evaluation requirement — we’ll help you assess how the data can support your use case.

Dr. Alexander Kel CEO And CSO

Prof. Dr. Alexander Kel

CEO & CSO, geneXplain GmbH · Co-author of TRANSFAC and Genome Enhancer

[email protected]

Explore the platform

Three ways in, one body of knowledge

TRANSFAC Workspace

Run the whole regulatory analysis yourself, on one no-code platform.

Open

TRANSFAC Expert

Give us the target or disease; our scientists deliver the answer.

Open

Overview

See all three ways in and pick the one that fits your team.

Open

Ground your models on curated biology

License the TRANSFAC Knowledge Graph, or request the schema to see how the three layers connect.