Background
Since 2024, I am an Associate Professor at the University of Modena and Reggio Emilia, where I work on Deep Learning, Vision-and-Language integration, Large-Scale models and Multimedia. I teach in the courses of "Computer Vision and Cognitive Systems," Scalable AI, and Computer Architecture. My research interests span various areas, including Vision-and-Language integration, Multimodal Retrieval, Image and Video Captioning, Visual-Semantic alignment, Large-Scale model development, HPC and Embodied AI.
I have authored more than 180 publications in international journals and conferences. Currently, I serve as an Associate Editor for Computer Vision and Image Understanding, Pattern Recognition and Pattern Recognition Letters, and act as an Area Chair for ICCV, ECCV, AAAI and major multimedia conferences. I am also a Scholar in the ELLIS society (European Laboratory for Learning and Intelligent Systems), where I coordinate the Modena ELLIS Unit.
From 2020 to 2026, I held the position of deputy director at the Interdepartmental Center on Digital Humanities at the University of Modena and Reggio Emilia. Earlier in my career, in 2017, I worked at the Facebook AI Research laboratory in Paris under the supervision of Hervé Jégou. During that time, I worked on the development of a video-matching algorithm that was adopted in production on the social network to detect abusive content.
News
We have open research collaborator and post-doc positions within the MINERVA and ELLIOT EU projects. If you are interested, please get in touch!
Program Chair for ICIAP 2027
Happy to share that I will serve as a Program Chair for ICIAP 2027, the biennial conference of CVPL (the Italian Association for Research in Computer Vision, Pattern Recognition and Machine Learning). The conference will be hosted in Modena on 13-17 September 2027, with workshops and tutorials at the Engineering Campus and the main conference at the Accademia Militare di Modena.
Paper accepted at TPAMI
Our paper "Recurrence Meets Transformers for Universal Multimodal Retrieval" has been accepted to TPAMI!
Paper accepted at EMNLP 2026
Our paper "CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models" has been accepted to EMNLP 2026!
Area Chair for AAAI 2027
Happy to share that I will serve as an Area Chair for AAAI 2027!
Leading WP2 of the ELLIOT project
I will lead WP2 “Foundation Models” of ELLIOT (European Lighthouse on Linked Open and Trustworthy Multimodal Foundation Models), a Horizon Europe project bringing together the main European actors working on multimodal foundation models.
Paper accepted at ECCV 2026
Our paper "A Scalable Vector Graphics Latent Space" has been accepted to ECCV 2026!
MSCA Doctoral Network grant for the ELLE project
Happy to share that our lab is a partner in ELLE (European Doctoral Network on Secure and Safe Artificial Intelligence in the Age of Foundation Models), a Marie Skłodowska-Curie Doctoral Network funded by the European Commission. ELLE will train 15 doctoral researchers across European institutions over 48 months, advancing research on AI safety and foundation models — directly aligned with our research agenda. 👉 Read the full article on the UNIMORE magazine
Technology alliance with AMD Silo AI on Physical AI
Our lab has entered a multi-year technology alliance with AMD Silo AI on Physical AI and open-source Vision-Language-Action stacks for AMD Instinct GPUs, funding two PhD positions. 👉 Read the announcement on the AMD blog
Paper accepted at CVPR 2026!
Our paper "ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering" has been accepted to CVPR 2026!
Featured publications
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
Alberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi, Davide Caffagni, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
A Scalable Vector Graphics Latent Space
Leonardo Zini, Elia Frigieri, Lorenzo Baraldi
Proceedings of the 19th European Conference on Computer Vision
vHector and HeisenVec: Scalable Vector Graphics Generation Through Large Language Models
Leonardo Zini, Elia Frigieri, Sebastiano Aloscari, Lorenzo Baraldi
NeurIPS 2025, Datasets and Benchmarks track
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
Lorenzo Baraldi, Davide Bucciarelli, Federico Betti, Marcella Cornia, Lorenzo Baraldi, Nicu Sebe, Rita Cucchiara
ICCV 2025
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
Luca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi, Fabrizio Falchi, Rita Cucchiara
ICCV 2025
MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
Vittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia, Lorenzo Baraldi, Costantino Grana, Rita Cucchiara, Elisa Ficarra
ICCV 2025
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
CVPR 2025
Hyperbolic Safety-Aware Vision-Language Models
Tobia Poppi, Tejaswi Kasarla, Pascal Mettes, Lorenzo Baraldi, Rita Cucchiara
CVPR 2025 Highlight
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
CVPR 2025
Causal Graphical Models for Vision-Language Compositional Understanding
Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi, Rita Cucchiara
ICLR 2025
Courses
Scalable AI (2025/2026)
AI Engineering
Lorenzo Baraldi, Giuseppe Fiameni, Sara Sarto
Computer Vision and Cognitive Systems (2025/2026)
AI Engineering
Lorenzo Baraldi, Vittorio Cuculo
Architettura dei Calcolatori (2025/2026)
Course material
· Upcoming exams
Architettura dei Calcolatori
Rita Cucchiara, Lorenzo Baraldi