Floris van der Flier
florisvdf.bsky.social
Floris van der Flier
@florisvdf.bsky.social
PhD Student @ WUR Bioinformatics,
ML & Protein Engineering
https://github.com/florisvdf
Thanks for the kind words Rens 😁
June 29, 2026 at 2:59 PM
In summary
- Uncertainty-aware acquisition can yield better identification of hit variants
- Favoring increased model uncertainty is preferred for single mutants
- Matching objectives to downstream use cases is worth exploring, even when traditional approaches dominate on average
June 29, 2026 at 10:22 AM
This matters, because practitioners don't need general-purpose models, they need models that excel on their dataset
June 29, 2026 at 10:22 AM
Key findings:

Traditional regression models still outperform preferential models on most datasets
But: Uncertainty-aware acquisition functions consistently win, especially when acquiring in high-uncertainty regions
Preferential models show clear advantages on specific datasets
June 29, 2026 at 10:21 AM
We built a retrospective evaluation protocol using quantile cross-validation (holding out high-value variants) and a custom recovery metric (measuring how well models + acquisition functions prioritize those held-out variants).
June 29, 2026 at 10:21 AM
We tailored model development and evaluation to match the real objective of acquisition. We tested two techniques:

Preferential learning → mirror the goal of variant selection
Uncertainty quantification → provide information about model confidence to guide acquisition
June 29, 2026 at 10:21 AM
On top of that, models are often inaccurate on new data, yet our acquisition strategies ignore where models fail. Model uncertainty can reveal these failure modes and guide smarter variant selection.
June 29, 2026 at 10:21 AM
Protein engineering models are built to predict exact properties, but when we use them, we only care if a variant is better than what we have. We evaluate them on ranking every variant in a dataset, but we only need them to find the rare, high-performing ones.
June 29, 2026 at 10:21 AM
We saw similar effects in our study, even for a combinatorial dataset we specifically designed to relate structural characteristics of mutations to their predictability. An augmented Potts model and a PLM based model achieved only marginal gains:
doi.org/10.1016/j.cs...
Redirecting
doi.org
March 13, 2026 at 11:17 AM
Are you using it as your main ide?
August 15, 2025 at 9:00 AM
(7/8) Despite the outcome, this was still a very fun project to work on. Many thanks my supervisors, Henning Redestig and Dick de Ridder, for their excellent guidance, and Luis-Cascao Pereira and David Estell for inspiring us to explore electrostatic quantities of proteins.
August 14, 2025 at 10:43 AM
(6/8) We suspect the limitations arise from imperfect computational tools, missing biological context (e.g., post-translational modifications, molecular crowding), and treating proteins as static objects. Still, we believe these results could help narrow the search space for future VEP strategies.
August 14, 2025 at 10:42 AM
(5/8) We trained and evaluated ChargeNet on datasets from ProteinGym, comparing it to evolutionary models both standalone and in ensemble. Across all tests, ChargeNet offered no advantage—suggesting its physical representations don’t add information beyond what evolutionary models already encode.
August 14, 2025 at 10:42 AM
(4/8) We tried to design a representation that is evolution agnostic and instead relies on the physical properties. This led to ChargeNet—a pipeline that predicts a variant’s structure with FoldX, computes its 3D electrostatic profile with APBS, and feeds that into a 3D CNN for property prediction.
August 14, 2025 at 10:42 AM
(3/8) SotA VEP models are evolutionary and excel at spotting “unnatural” mutations, which often correlate with loss of function. However, by pretraining only on natural sequences, such models may miss out on mutations that wouldn’t occur in nature, but still lead to interesting changes in function.
August 14, 2025 at 10:41 AM
(2/8) Unfortunately, our strategy did not yield the results we hoped for; our model is outperformed by existing models in every scenario we tested. Nevertheless, we find it important to share our findings so that others can build on them, refine their approach, or take more promising directions.
August 14, 2025 at 10:41 AM
Question: Could the experimental uncertainty arise from the dynamics of the protein in solution? And if so, are there examples where this uncertainty strongly correlates with dynamics?
February 3, 2025 at 3:52 PM
Once again looking amazing. Are you considering at some point curating a gallery to showcase what's possible with molecular nodes? Or does this already exist and did I miss it? 🫣
December 6, 2024 at 9:21 AM