Patrick C H Ma, Keith C C Chan
Journal: Journal of computational biology : a journal of computational molecular cell biology 2008;15(4):431-43
PMID: 18435571
To classify proteins into functional families based on their primary sequences, popular algorithms such as the k-NN-, HMM-, and SVM-based algorithms are often used. For many of these algorithms to perform their tasks, protein sequences need to be properly aligned first. Since the alignment process can be error-prone, protein classification may not be performed very accurately. To improve classification accuracy, we propose an algorithm, called the Unaligned Protein SEquence Classifier (UPSEC), which can perform its tasks without sequence alignment. UPSEC makes use of a probabilistic measure to identify residues that are useful for classification in both positive and negative training samples, and can handle multi-class classification with a single classifier and a single pass through the training data. UPSEC has been tested with real protein data sets. Experimental results show that UPSEC can effectively classify unaligned protein sequences into their corresponding functional families, and the patterns it discovers during the training process can be biologically meaningful.
Full Text Sources:
© Copyright 2026, Nutrition Evidence
We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.