The aim of this endeavor is to explore the application of multivariate statistical methods in predicting whether the patient has kidney stone disease (KSD) or not. A public dataset has been created on Kaggle containing information on patients with KSD, including age, gender, and stone type, as well as laboratory test results. A systematic methodology has been followed, starting with exploratory data analysis (EDA) to identify and understand the patterns and trends in the dataset. As a second step, principal component analysis (PCA) has been applied. Then, based on the PCA results, a logistic regression (LR) model has been built. As a last step, clustering algorithms, including k-nearest neighbors and hierarchical clustering, were utilized to categorize the patients and run a comparison. The results showed that logistic regression was accurate in predicting the development of KSD with high precision, and the results of the clustering steps were two clusters. This research revealed the predictive potential of multivariate statistical methods in detecting KSD, and they have major consequences regarding the diagnosis and treatment of KSD patients.
Exploring Multivariate Techniques for Analyzing Kidney Stone Risk Factors: Logistic Regression and Cluster Analysis Approach
1 view
1 Downloads