🇮🇹 Paper accepted at EVALITA 2026: INDAQA 2
-
Our paper, INDAQA2 - A Large Italian Narrative QA Benchmark: A CALAMITA 2026 Challenge, has been accepted at EVALITA Workshop 2026!
👏 Huge thanks to my co-authors Luca Moroni, Alberte Fernández-Castro, Elena Marafatto, Giacomo Garufi and Roberto Navigli.
Abstract
Long-context comprehension and reasoning remain largely underexplored in the evaluation of Italian Large Language Models (LLMs). Existing Italian benchmarks primarily focus on short or medium-length inputs, offering limited insight into models’ ability to process extended narratives. To address this gap, we introduce INDAQA2, a substantially revised and expanded version of INDAQA, a benchmark for narrative question answering on original Italian literary texts. The new version comprises an expanded corpus of 461 total books, introduces a multiple-choice question answering format alongside the original open-ended tasks, and features manually curated texts drawn exclusively from works originally written in Italian, thus avoiding artifacts introduced by translation. The benchmark evaluates long-context understanding over complete books of up to 250K tokens, testing complementary comprehension skills through a dual-structure design: global narrative understanding, assessed via questions derived from book summaries, and local precision, assessed via questions grounded in specific passages and entity-level details. By supporting both open-ended and multiple-choice question answering formats, INDAQA2 enables evaluation of both generative capabilities and discriminative reasoning, facilitating comprehensive and scalable comparison across models. Our evaluation of several Italian-specialized and multilingual models reveals significant performance disparities across task formats and highlights limitations in how current Italian models utilize extended contexts.
See you in Bari! 🇮🇹
INDAQA2 - A Large Italian Narrative QA Benchmark: A CALAMITA 2026 Challenge
INDAQA2 is a manually validated resource designed to fill an important gap in the evaluation of narrative understanding for LLMs on Italian narrative texts. It spans over 460 books and includes more than 25k question-answering pairs. The benchmark features a dual evaluation procedure: a subset of open-ended questions is also presented in multiple-choice format, enabling a more reliable assessment of structural biases introduced by evaluation and prompting formats.
Citations
If you found our work valuable, please cite it as:
-
Plain text
Luca Gioffré, Luca Moroni, Alberte Fernández-Castro, Elena Marafatto, Giacomo Garufi, and Roberto Navigli. 2026. INDAQA2 - A Large Italian Narrative QA Benchmark: A CALAMITA 2026 Challenge. In Proceedings of the 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026), Bari, Italy. CEUR Workshop Proceedings.Markdown
[INDAQA2 - A Large Italian Narrative QA Benchmark: A CALAMITA 2026 Challenge](https://apa.dipsco.unitn.it/evalita2026/69.pdf) (Gioffré et al., EVALITA 2026)BibTeX
@inproceedings{gioffre-etal-2026-INDAQA2, title = {INDAQA2 - A Large Italian Narrative QA Benchmark: A CALAMITA 2026 Challenge}, author = {Gioffr{\'e}, Luca and Moroni, Luca and Fernández-Castro, Alberte and Marafatto, Elena and Garufi, Giacomo and Navigli, Roberto}, editor = {Cutugno, Francesco and Miaschi, Alessio and Aprosio, Alessio Palmero and Rambelli, Giulia and Siciliani, Lucia and Stranisci, Marco Antonio}, booktitle = { Proceedings of the 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026)}, month = feb, year = {2026}, address = {Bari, Italy}, publisher = {CEUR Workshop Proceedings}, url = {https://apa.dipsco.unitn.it/evalita2026/69.pdf} }