A comparative study of supervised and unsupervised machine learning algorithms applied to human microbiome.

K Dhuli, A Macchia, G Bonetti, K Donato, M Bertelli, E Kalluçi, B Preni, X Dhamo, E Noka, S Bardhi, L J M Zambrano, S Janaqi

Journal: La Clinica terapeutica 2024;175(3):98-116

PMID: 38767067

Abstract

BACKGROUND

The human microbiome, consisting of diverse bacte-rial, fungal, protozoan and viral species, exerts a profound influence on various physiological processes and disease susceptibility. However, the complexity of microbiome data has presented significant challenges in the analysis and interpretation of these intricate datasets, leading to the development of specialized software that employs machine learning algorithms for these aims.

METHODS

In this paper, we analyze raw data taken from 16S rRNA gene sequencing from three studies, including stool samples from healthy control, patients with adenoma, and patients with colorectal cancer. Firstly, we use network-based methods to reduce dimensions of the dataset and consider only the most important features. In addition, we employ supervised machine learning algorithms to make prediction.

RESULTS

Results show that graph-based techniques reduces dimen-sion from 255 up to 78 features with modularity score 0.73 based on different centrality measures. On the other hand, projection methods (non-negative matrix factorization and principal component analysis) reduce dimensions to 7 features. Furthermore, we apply supervised machine learning algorithms on the most important features obtained from centrality measures and on the ones obtained from projection methods, founding that the evaluation metrics have approximately the same scores when applying the algorithms on the entire dataset, on 78 feature and on 7 features.

CONCLUSIONS

This study demonstrates the efficacy of graph-based and projection methods in the interpretation for 16S rRNA gene sequencing data. Supervised machine learning on refined features from both approaches yields comparable predictive performance, emphasizing specific microbial features-bacteroides, prevotella, fusobacterium, lysinibacillus, blautia, sphingomonas, and faecalibacterium-as key in predicting patient conditions from raw data.

Address: Department of Applied Mathematics, Faculty of Natural Sciences, University of Tirana, Tirana, Albania.; Department of Mathemat-ics, Faculty of Engineering Mathematics and Engineering Physics, Polytechnic University of Tirana, Tirana, Albania.; Department of Applied Statistics and Informatics, University of Tirana, Tirana, Albania.; MAGI'S LAB, Rovereto (TN), Italy.; MAGI'S LAB, Rovereto (TN), Italy.; Department of Pharmaceutical Sciences, University of Perugia, Perugia, Italy.; MAGI EUREGIO, Bolzano, Italy.; MAGISNAT, Atlanta Tech Park, Peachtree Corners, GA, USA.; MAGI'S LAB, Rovereto (TN), Italy.; MAGI EUREGIO, Bolzano, Italy.; MAGISNAT, Atlanta Tech Park, Peachtree Corners, GA, USA.; Computational Biology Group, Precision Nutrition and Cancer Research Program, IMDEA Food Institute, Madrid, Spain.; EuroMov Digital Health in Motion, University of Montpellier IMT Mines Ales, France.

Link outs

Bant logo

© Copyright 2026, Nutrition Evidence

NED wishes to thank the following organisations for their support:

We use cookies to improve your experience and analyze site traffic with Google Analytics. By continuing to use our site, you agree to our use of cookies. Learn more.