🇮🇹 Paper accepted at CLiC-it 2026: Dromedario 3 🐪
-
Our paper, Dromedario 3: Localizing the Tülu 3 Dataset to Italian, has been accepted at CLiC-it Conference 2026!
👏 Huge thanks to my co-authors Marina Iuliana Aur, Francesco Ortame, Luca Moroni, Alberte Fernández-Castro, Elena Marafatto and Roberto Navigli.
See you in Palermo! 🇮🇹
Dromedario 3
Dromedario 3 is a large-scale Italian instruction-tuning dataset built from the Tülu 3 SFT mixture through a documented pipeline of taxonomy-based filtering, translation, and post-translation quality control. We validate Dromedario 3 training two models, Minerva 7B and Llama 3 8B on different recipes, and we found that Italian SFT improves the performance of models with stronger Italian pretraining of up to 30% compared to the original Tülu 3 mixture.
This result suggests that language-specific SFT data matters; in this context, translated English data of silver quality is a first step toward building more capable Italian models.
Citations
If you found our work valuable, please cite it as:
-
Plain text
Luca Gioffré, Marina Iuliana Aur, Francesco Ortame, Luca Moroni, Alberte Fernández-Castro, Elena Marafatto, and Roberto Navigli. 2026. Dromedario 3: Localizing the Tülu 3 Dataset to Italian. In Proceedings of the Twelfth Italian Conference on Computational Linguistics (CLiC-it 2026), Palermo, Italy.Markdown
Dromedario 3: Localizing the Tülu 3 Dataset to Italian (Gioffré et al., CLiC-it 2026)BibTeX
@inproceedings{gioffre-2026-dromedario3, title = {Dromedario 3: Localizing the T{\"u}lu 3 Dataset to Italian}, author = {Gioffr{\'e}, Luca and Aur, Marina Iuliana and Ortame, Francesco and Moroni, Luca and Fern{\'a}ndez-Castro, Alberte and Marafatto, Elena and Navigli, Roberto}, booktitle = {Proceedings of the Twelfth Italian Conference on Computational Linguistics (CLiC-it 2026)}, editor = {Basile, Valerio and Croce, Danilo and Passaro, Lucia C. and Pirrone, Roberto}, year = {2026}, address = {Palermo, Italy}, month = sep }