Chloe Hsu, Hunter Nisonoff, Clara Fannjiang, Jennifer Listgarten
Journal: Nature biotechnology 2022;40(7):1114-1122
PMID: 35039677
Machine learning-based models of protein fitness typically learn from either unlabeled, evolutionarily related sequences or variant sequences with experimentally measured labels. For regimes where only limited experimental data are available, recent work has suggested methods for combining both sources of information. Toward that goal, we propose a simple combination approach that is competitive with, and on average outperforms more sophisticated methods. Our approach uses ridge regression on site-specific amino acid features combined with one probability density feature from modeling the evolutionary data. Within this approach, we find that a variational autoencoder-based probability density model showed the best overall performance, although any evolutionary density model can be used. Moreover, our analysis highlights the importance of systematic evaluations and sufficient baselines.
© 2022. The Author(s), under exclusive licence to Springer Nature America, Inc.
Other Literature Sources:
Full Text Sources:
© Copyright 2026, Nutrition Evidence
We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.