Large language models (LLMs) have shown promising capabilities in assisting researchers and developers in different fields of cybersecurity. This work investigates whether 11 state-of-the-art LLMs can be used for source code vulnerability analysis across three different use cases and four publicly available benchmark datasets. More specifically, we examined Android, smart contract and IoT source code, containing vulnerabilities from Open Worldwide Application Security Project (OWASP) Mobile Top 10, Common Weakness Enumeration (CWE) databases, and smart contract related vulnerabilities. Moreover, we explored whether LLMs could detect potentially privacy-invasive actions and if retrieval-augmented generation (RAG) could improve the performance of LLMs in vulnerability detection. Our results reveal that no single LLM is consistently better-performing compared to others across all use cases and datasets, whereas different models are the best performers in different use cases and datasets. Thus, a careful LLM selection is necessary based on the unique characteristics of each use case.
KOULIARIDIS Vasileios;
KAROPOULOS Georgios;
KAMBOURAKIS Georgios;
2026-08-26
INDERSCIENCE PUBLISHERS
JRC143456
1753-0571 (online),
https://www.inderscience.com/info/inarticle.php?artid=154618,
https://publications.jrc.ec.europa.eu/repository/handle/JRC143456,
10.1504/IJACT.2026.154618 (online),
| Name | Country | City | Type |
|---|
This document is only visible at the Commission level.
You are not authorized to publish or distribute it outside the European Commission.
This is a public document. You can share this publication.
Datasets
| ID | Title | Public URL |
|---|
Dataset collections
| ID | Acronym | Title | Public URL |
|---|
Scripts / source codes
| Description | Public URL |
|---|
Additional supporting files
| File name | Description | File type |
|---|