Within the scope of Philadelphia chromosome-negative myeloproliferative neoplasms, the effective differentiation of Primary Myelofibrosis from Essential Thrombocythaemia and Polycythaemia Vera is of vital importance for prognosis estimation and treatment approach planning. This study aimed to distinguish Primary Myelofibrosis samples from the Essential Thrombocythaemia - Polycythaemia Vera group using two gene expression datasets obtained from the Gene Expression Omnibus repository. A dataset of 69 samples for training and data from 33 untreated samples for independent external validation were converted to the gene level via GPL570 annotation and matched in a common feature space. Data preprocessing, feature selection, and modelling procedures were designed within a pipeline to minimize the risk of data leakage. Hyperparameter optimization was performed using the GridSearchCV technique with iterative stratified cross-validation. Similar Average Precision and Balanced Accuracy values were obtained during the validation phase performed on the training data. In the independent external validation phase, the Linear Support Vector Machine method stood out with higher Average Precision and Balanced Accuracy values, while the Elastic Net Logistic Regression method provided higher values in terms of the Receiver Operating Characteristic - Area Under Curve metric. The findings reveal that a pipeline design using open-source data from the Gene Expression Omnibus archive can be used to distinguish Primary Myelofibrosis samples.
Keywords
Primary Myelofibrosis, Gene Expression, Linear SVM, Elastic Net Logistic Regression, and PR-AUC.