19,443 Complexes · 48 GBPDBbind-cn Benchmark (v2020/v2024)
The standard curated collection of experimentally measured binding affinity data (Kd, Ki, IC50) for biomolecular complexes in the Protein Data Bank.
Standardized biological training corpora, blind benchmark evaluations, and multi-billion compound screening libraries powering modern generative biology.
19,443 Complexes · 48 GBThe standard curated collection of experimentally measured binding affinity data (Kd, Ki, IC50) for biomolecular complexes in the Protein Data Bank.
120 Blind TargetsThe gold-standard biennial Critical Assessment of Structure Prediction targets evaluating blind single-chain, multimer, and protein-ligand predictions.
700k+ Molecules across 17 DatasetsComprehensive benchmark suite for molecular machine learning across quantum mechanics, physical chemistry, biophysics, and physiology.
42M+ Cells across 35 OrgansGlobal collaborative effort mapping all human cell types with single-cell RNA sequencing, ATAC-seq, and spatial transcriptomics.
65M Families · 2.2 Billion Sequences (560 GB)Massive reference database combining UniProt and environmental metagenomic assemblies, foundational for deep MSA searches in AlphaFold and OpenFold.
5.5 Billion Synthesizable CompoundsThe largest commercially accessible virtual screening library for ultra-large library docking and generative lead identification.