Deep-learning co-folding models are becoming increasingly adopted in drug discovery. However, benchmark results accompanying model releases can be hard to replicate, access, or inspect. CoFold Arena aims to provide a centralized, maximally transparent, and regularly updated benchmark for co-folding models on therapeutically relevant tasks.
Compared to existing benchmarks (e.g. FoldBench, PXMeter), CoFold Arena has three differentiating aims:
Availability and transparency. All model predictions upstream of headline metrics can be inspected in-browser or easily downloaded for reanalysis.
Up-to-date with the current PDB. All models are run on weekly PDB releases to provide a maximally recent, common, and evolving benchmark set. This (1) enables head-to-head comparisons of all available models instead of excluding models with later training cutoffs, and (2) eliminates the possibility of bias in retrospective evaluations by model developers.
Customization. Instead of rigidly defined test sets, arbitrary PDB release ranges can be selected in-browser for evaluation.
Master cutoff: Evaluations are available for structures released after September 9, 2024, past the self-reported training cutoff of all models (except Protenix-v1-20250630).
Below we describe the methodology for the antibody-antigen co-folding leaderboard. This page will be updated with methodology for other leaderboards as they become available.
Models are evaluated on SAbDab entries meeting all of the following, applied in order:
SABDAB_ID), taking its earliest-released structure.auth_asym_ids in the deposited mmCIF are skipped.X tokens, so every model sees the same query sequence.0.8·ipTM + 0.2·pTM). AF-Multimer runs the v3 weights (model_1_multimer_v3) and produces a single structure.Each prediction is compared to the trimmed native with DockQ over the antibody–antigen interface(s) using tinyprot. The leaderboard summarizes the per-target DockQ two ways:
Turning on Show uncertainty adds a paired bootstrap over the target panel: targets are resampled with replacement and every model is re-scored on the same resample, giving each metric a 95% confidence interval and each model a rank range.