Σύγχρονες μέθοδοι ανάλυσης δεδομένων επιβίωσης υψηλής διάστασης : επιλογή μεταβλητών και μοντέλα πρόβλεψης
Modern methods for high - dimensional survival data analysis : variable selection and predictive models
View/ Open
Keywords
Ανάλυση επιβίωσης ; Επιλογή μεταβλητών ; Μοντέλα πρόβλεψης ; Ποινικοποίηση ; Δεδομένα υψηλής διάστασης ; Μηχανική μάθησηAbstract
The analysis of high – dimensional survival data constitutes a particularly challenging statistical problem, as the number of available explanatory variables can be far greater than the sample size, while the presence of censored observations must also be taken into account. In such settings, selecting variables with substantial prognostic information and developing models with the best possible generalization capability are key challenges. This thesis presents the fundamental concepts and classical models of Survival Analysis, and subsequently examines modern methods for high – dimensional data.
In high – dimensional environments (𝑝≫𝑛), where overfitting and multicollinearity limit predictive performance, a two – stage strategy is proposed. This approach initially reduces dimensionality through feature screening (e.g., RMST), followed by the application of penalized regression (regularization) such as Broken Adaptive Ridge, Improved Adaptive Lasso, for simultaneous parameter estimation and variable selection. Modern techniques incorporating internal regularization mechanisms that produce sparse and interpretable results in high – dimensional survival data are also examined. Furthermore, advanced Machine Learning techniques are considered for detecting complex nonlinear relationships, including Survival Trees (OST, OSST), Oblique Random Survival Forests (Oblique RSF), Boosting algorithms, Survival SVM and approaches based on Kaplan – Meier weighted schemes.
The thesis concludes with the application of selected techniques to a real high – dimensional survival dataset, aiming to develop an optimal predictive model based on the Concordance Index (C-index) and identifies the most important genes associated with survival time.
Overall, this work highlights the importance of appropriate variable selection, proper validation procedures, and the combined utilization of statistical and computational methods to develop reliable prognostic models for high – dimensional survival data.


