Wikidata has emerged as a major open repository of scholarly metadata, yet its characteristics and limitations as a source for diversity analyses are not well documented. We present DivinWD (DIVersity IN WikiData), a curated dataset consisting of more than 23 million triples, produced via a fully open-source processing pipeline to enable the study of diversity in scientific publications represented in Wikidata. The dataset comprises over 1.2 million scholarly articles published between 2010 and 2024, enriched by integrating Wikidata with 5 external bibliographic sources--Crossref, Dimensions, OpenAlex, Scopus, and Semantic Scholar--and augmenting them using the Genderize API, to enrich metadata on language, field of study, authorship, gender, geographic origin, and institutional affiliation.
Our analysis documents systematic coverage biases and infrastructural artifacts affecting Wikidata's scholarly content, highlighting important considerations for reuse. By releasing the dataset and pipeline, this work provides a transparent foundation for future research on diversity in science and for the development and evaluation of open, reproducible bibliometric indicators.
SALETTI Zeno;
CONSONNI Cristian;
FRAU AMAR Pedro;
GOMEZ Emilia;
2026-07-20
Association for the Advancement of Artificial Intelligence (AAAI)
JRC145784
https://ojs.aaai.org/index.php/ICWSM/article/view/42790,
https://publications.jrc.ec.europa.eu/repository/handle/JRC145784,
10.1609/icwsm.v20i1.42790 (online),
| Name | Country | City | Type |
|---|
This document is only visible at the Commission level.
You are not authorized to publish or distribute it outside the European Commission.
This is a public document. You can share this publication.
Datasets
| ID | Title | Public URL |
|---|
Dataset collections
| ID | Acronym | Title | Public URL |
|---|
Scripts / source codes
| Description | Public URL |
|---|
Additional supporting files
| File name | Description | File type |
|---|