Ensuring confidentiality and personal data protection in large language models
Διασφάλιση της εμπιστευτικότητας και προστασίας προσωπικών δεδομένων στα μεγάλα γλωσσικά μοντέλα

Master Thesis
Author
Atherinis - Spartiotis, Leandros
Αθερίνης - Σπαρτιώτης, Λέανδρος
Date
2026-05Advisor
Petasis, GeorgiosΠετάσης, Γεώργιος
View/ Open
Abstract
This thesis deals with how Large Language Models can protect confidentiality and personal data. It starts from the observation that existing solutions, meaning legislation and infrastructure security measures, are not enough since the model itself can reveal sensitive information through conversation. Protection therefore has to become part of the model’s own behavior. After covering the legal framework, the risks and existing techniques, th thesis tests two training approaches: first, training a model in the role of the defender, acting as a customer service representative, with reinforcement learning through a PPO trainer. Then, a new framework proposed in this work, the LLM-GAN, where an attacker model and the defender are trained against each other. The main conclusion is that PPO training on its own has weaknesses, whereas with the LLM-GAN, where the defender constantly faces an opponent that adapts in order to extract information from it and tries out many strategies, the model receives more meaningful training and is more secure across a wider range of attacks and situations.


