Clustered tree regression to learn protein energy change with mutated amino acid.

Hongwei Tu, Yanqiang Han, Zhilong Wang, Jinjin Li

Journal: Briefings in bioinformatics 2022;23(6):bbac374

PMID: 36124753

Abstract

Accurate and effective prediction of mutation-induced protein energy change remains a great challenge and of great interest in computational biology. However, high resource consumption and insufficient structural information of proteins severely limit the experimental techniques and structure-based prediction methods. Here, we design a structure-independent protocol to accurately and effectively predict the mutation-induced protein folding free energy change with only sequence, physicochemical and evolutionary features. The proposed clustered tree regression protocol is capable of effectively exploiting the inherent data patterns by integrating unsupervised feature clustering by K-means and supervised tree regression using XGBoost, and thus enabling fast and accurate protein predictions with different mutations, with an average Pearson correlation coefficient of 0.83 and an average root-mean-square error of 0.94kcal/mol. The proposed sequence-based method not only eliminates the dependence on protein structures, but also has potential applications in protein predictions with rare structural information.

© The Author(s) 2022. Published by Oxford University Press. All rights reserved. For Permissions, please email: [email protected].

Address: Key Laboratory of Thin Film and Microfabrication of Ministry of Education, Department of Micro/Nano Electronics, Shanghai Jiao Tong University, Shanghai, 200240, China.
Bant logo

© Copyright 2026, Nutrition Evidence

NED wishes to thank the following organisations for their support:

We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.