Μέτρα απόστασης και ομοιότητας σε πολυδιάστατους χώρους και η χρήση τους σε προβλήματα ομαδοποίησης πολυδιάστατων δεδομένων
Distance and similarity measures in multidimensional spaces and their use in multidimensional data clustering problems

View/ Open
Keywords
Μέτρα απόστασης ; Μέτρα ομοιότητας ; Ανάλυση κατά συστάδες ; Πολυδιάστατα δεδομένα ; Πολυμεταβλητή στατιστική ανάλυση ; Ομαδοποίηση δεδομένωνAbstract
Cluster Analysis is one of the most important branches of Statistics and Data Science, with applications across a wide range of scientific and technological fields. Its primary objective is to discover underlying structures or patterns within datasets by grouping together observations that exhibit similarities while separating those that display significant differences. A central concept in this process is the notion of distance or similarity, as the measure selected to quantify the proximity between objects directly determines the outcome and quality of the clustering process.
The choice of different distance metrics may lead to substantially different clustering results, even when the same clustering algorithm is applied to the same dataset. The literature presents a wide variety of alternative metrics, ranging from classical Euclidean distances to more specialized measures such as the Mahalanobis distance and similarity indices such as the Jaccard and Cosine measures. Each metric possesses distinct theoretical advantages and limitations, as well as different areas of applicability, making their study and comparison essential for understanding their behavior in high-dimensional data spaces.
Within this context, the present master’s thesis focuses on three main objectives. First, it provides a systematic presentation of the most widely used distance metrics and similarity measures found in the literature, including their mathematical properties and the theoretical relationships that connect them. Second, it investigates the relationship between distance and similarity, as well as the methodologies for transforming one type of measure into the other, with particular emphasis on applications in machine learning and data mining techniques. Third, it experimentally evaluates different distance metrics by applying clustering algorithms to multidimensional datasets to assess their effectiveness under realistic scenarios.
The combined consideration of theoretical foundations and practical applications is expected to provide a comprehensive understanding of the role that the choice of a distance metric plays in the clustering process. In this way, the thesis aims not only to highlight the theoretical aspects of the problem but also to propose practical guidelines for the appropriate selection of distance metrics according to the characteristics of the data and the specific requirements of each application.


