Although automatic sentiment analysis has been widely studied in the past decade, multilingualism remains an issue impairing real-life applications. This paper describes the development and evaluation of XLM-RLnews-8, a model based on XLM-RoBERTa-Large, domain-adapted using a novel dataset of multilingual news articles and fine-tuned for the tripartite sentiment analysis task. In addition, it provides a quantitative analysis of the Unified Multilingual Sentiment Analysis Benchmark, the dataset used for the fine-tuning and the in-domain evaluation. The model is also out-of-domain evaluated, on the IMDb dataset and a new multilingual news headlines silver dataset, and its performance is compared with current State-of-the-Art multilingual models.
DI NUOVO Elisa;
CARTIER Emmanuel;
DE LONGUEVILLE Bertrand;
2025-07-23
SPRINGER-VERLAG BERLIN
JRC137542
1611-3349 (online),
https://link.springer.com/chapter/10.1007/978-3-031-70242-6_3,
https://publications.jrc.ec.europa.eu/repository/handle/JRC137542,
10.1007/978-3-031-70242-6_3 (online),
| Name | Country | City | Type |
|---|
This document is only visible at the Commission level.
You are not authorized to publish or distribute it outside the European Commission.
This is a public document. You can share this publication.
Datasets
| ID | Title | Public URL |
|---|
Dataset collections
| ID | Acronym | Title | Public URL |
|---|
Scripts / source codes
| Description | Public URL |
|---|
Additional supporting files
| File name | Description | File type |
|---|