Etowah Adams
etowah0.bsky.social
Etowah Adams
@etowah0.bsky.social
enjoying and bemoaning biology. phd student
@columbia prev. @harvardmed @ginkgo @yale
Additional thanks to @moalquraishi.bsky.social, Fergus Imrie, Frank von Delft, @jchodera.bsky.social, Charlotte Deane, and the people who believed in OpenBind to get it initially funded: Charlie Harris, Matt Clifford, Andrew Crawford, @tomwestgarth.bsky.social
August 21, 2026 at 6:15 PM
Huge thanks to the entire OpenBind team and its funders, @asapdiscovery.bsky.social , Syngenta, and @diamondlightsource.bsky.social , and to everyone who helped generate, process, and benchmark these data.
August 21, 2026 at 6:15 PM
We hope OB0 will be a useful model for the community to use and build on. As OpenBind generates more data, OB0 will serve as a baseline for measuring its impact and help guide what data we collect next.
August 21, 2026 at 6:15 PM
And these are hard drug targets! On Dengue and Zika RdRp, no cofolding model exceeded 10% top-25 success (best: 7.7% and 6.7%). The target systems combine flexible binding sites, multiple pockets, and poor training-set coverage.
August 21, 2026 at 6:15 PM
We’re also releasing 717 discovery-relevant ligand-bound structures across three targets: FatA, Dengue RdRp, and Zika RdRp. We believe that releasing hard, real-world evaluation sets alongside models is a good practice moving forward.
August 21, 2026 at 6:15 PM
Pre-training on new data may provide benefits beyond fine-tuning. OB0, which saw the fragment structures during pre-training, outperforms a version of OF3p2 fine-tuned on those fragments when predicting follow-on compounds.
August 21, 2026 at 6:15 PM
How does OB0 do in a fragment-based drug discovery context? Bound fragment structures guide follow-on design. We previously showed that fine-tuning on fragment structures improves follow-on predictions. What if they’re seen during pre-training?
August 21, 2026 at 6:15 PM
This highlights an important distinction. More data doesn't necessarily translate into improve models. Models improve most when new data fills gaps in the training set and provides relevant examples they didn’t have before.
August 21, 2026 at 6:15 PM
The biggest gains come where new data adds more relevant training examples: test complexes that become more similar to the 2025 training set see substantial improvements in accuracy. Those that stay in the same similarity bin show little difference.
August 21, 2026 at 6:15 PM
Do four more years of training data make the test set easier? Less than you might think. Most test complexes are about as similar to the 2021 training set as they are to the 2025 training set.
August 21, 2026 at 6:15 PM
OB0 is the first in a series of OpenBind models, with each release incorporating newer PDB structures and OpenBind data. It sets a baseline for measuring the impact of OpenBind data. Future models will add further new small-molecule capabilities.
August 21, 2026 at 6:15 PM
Unlike OpenFold3-preview2 and AF3, trained on data through 2021, OpenBind-0 (OB0) uses data through June 2025 and the forthcoming OpenFold3 architecture. It is competitive with leading cofolding models and adds chemical steering to improve ligand validity.
August 21, 2026 at 6:15 PM
Thank you for your help interpreting features :) It was really special having the community squint at these features with us. Lots of weird features means that everyone sees different/new things!
February 10, 2025 at 4:13 PM
Thanks to our coauthors Minji Lee, Steven Yu, and @moalquraishi.bsky.social!

Check work from Elana Pearl on using SAEs on pLMs too!
x.com/ElanaPearl/s...
February 10, 2025 at 4:12 PM
Such hard-to-interpret but predictive features could result from biases and limitations in our datasets and models. However, there is another intriguing possibility: these features may correspond to biological mechanisms that have yet to be discovered.
February 10, 2025 at 4:12 PM
Not all predictive features are easily interpretable. For instance, a predictor for membrane localization, latent L28/3154, predominantly activates on poly-alanine sequences, whose functional relevance remains unclear.
February 10, 2025 at 4:12 PM