We present polyNarrative, a new multilingual dataset of news articles, annotated for narratives. Narratives are overt or implicit claims, recurring across articles and languages, promoting a specific interpretation or viewpoint on an ongoing topic, often propagating mis/disinformation. We developed two-level taxonomies with coarse- and fine-grained narrative labels for two domains: (i) climate change and (ii) the military conflict between Ukraine and Russia. We collected news articles in four languages (Bulgarian, English, Portuguese, and Russian) related to the two domains and manually annotated them at the paragraph level. We make the dataset publicly available, along with experimental results of several strong baselines that assign narrative labels to news articles at the paragraph or the document level. We believe that this dataset will foster research in narrative detection and enable new research directions towards more multi-domain and highly granular narrative related tasks.
NIKOLAIDIS Nikolaos;
STEFANOVITCH Nicolas;
SILVANO Purificação;
DIMITROV Dimitar;
YANGARBER Roman;
GUIMARÃES Nuno;
SARTORI Elisa;
ANDROUTSOPOULOS Ion;
NAKOV Preslav;
DA SAN MARTINO Giovanni;
PISKORSKI Jakub;
2025-12-23
Association for Computational Linguistics
JRC141408
https://aclanthology.org/2025.acl-long.1513.pdf,
https://publications.jrc.ec.europa.eu/repository/handle/JRC141408,
10.18653/v1/2025.acl-long.1513 (online),
| Name | Country | City | Type |
|---|
This document is only visible at the Commission level.
You are not authorized to publish or distribute it outside the European Commission.
This is a public document. You can share this publication.
Datasets
| ID | Title | Public URL |
|---|
Dataset collections
| ID | Acronym | Title | Public URL |
|---|
Scripts / source codes
| Description | Public URL |
|---|
Additional supporting files
| File name | Description | File type |
|---|