Μελέτη της αποτελεσματικότητας των LLMs σε διαφορετικά επιστημονικά πεδία
Study of the effectiveness of LLMs across different scientific fields

View/ Open
Keywords
Large Language Models ; Evaluation ; Benchmarks ; Scientific fields ; Systematic literature reviewAbstract
This thesis investigates the effectiveness of Large Language Models (LLMs) across various
scientific disciplines, examining how their performance fluctuates depending on the specific
domain. As generative artificial intelligence is rapidly being integrated into both research and
professional environments, a critical evaluation of its capabilities and limitations is highly
necessary.
To this end, a Systematic Literature Review (SLR) was conducted, adhering to the PRISMA
guidelines and the PICOC framework. By analyzing 43 recent peer-reviewed studies, this
research maps the performance of LLMs across five core areas: STEM and Mathematics,
Computer Science and Programming, Medicine and Biosciences, Law, and Social Sciences
and Humanities.
The findings indicate that the effectiveness of LLMs is far from uniform. It heavily depends on
the degree of standardization within a field, the volume of available training data, and the
complexity of the required reasoning. While the models demonstrate exceptional performance
in highly structured tasks, such as generating isolated code snippets or solving basic
mathematical problems, they struggle significantly with multi-step logical reasoning and tasks
requiring nuanced clinical or legal judgment. Even when achieving high scores on standardized
benchmarks, they remain prone to hallucinations. Furthermore, a substantial performance gap
is evident when operating in low-resource languages or outside Anglo-centric cultural and
institutional contexts.
Ultimately, although optimization techniques—such as prompt engineering, fine-tuning, and
agentic workflows—drastically enhance model output, human oversight remains irreplaceable.
LLMs are best utilized as advanced assistive tools rather than autonomous decision-makers,
particularly in high-stakes environments.


