I am a PhD student at the Hasso Plattner Institute (HPI), working with
Prof. Gerard de Melo.
My research lies at the intersection of vision-language models, computer vision, and
vector graphics. I am interested in studying where current VLMs fall short in
visually and relationally complex settings, and in teaching models to generate and
reason with visual content in ways that more closely resemble how humans think and
communicate.
I am currently working with
Prof. Yael Vinker
and Prof. Antonio Torralba (MIT CSAIL) on a project to enable sketching capabilities
in the model's reasoning process, with the goal of building systems that can explain
concepts visually and collaborate with students in a shared whiteboard space. I also
worked with Prof. Andres Sevtšuk (MIT DUSP) on a project to detect social
interactions in semantically complex urban scenes, with the goal of mapping which
areas of modern cities are more socially vibrant than others.
Latest News
I am spending three months at MIT CSAIL as a visiting researcher in Prof. Antonio Torralba’s lab.
Our work
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
was accepted at AAAI 2026.
Our project
Whiteboard-literate AI
won a grant from the HPI–MIT Designing for Sustainability program.
Our work
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision
was accepted at ICML 2025.
Publications
2026
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
Liu Liu, Alexandra Schild, Marco Cipriano, Fatimeh Al Ghannam,
Freya Tan, Gerard de Melo, and Andres Sevtsuk.
Proceedings of the AAAI Conference on Artificial Intelligence, 40(45),
38935-38942.
2025
Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision
Marco Cipriano, Moritz Feuerpfeil, and Gerard de Melo.
Forty-second International Conference on Machine Learning.
2024
ELSA: Evaluating Localization of Social Activities in Urban Streets using
Open-Vocabulary Detection
Maryam Hosseini, Marco Cipriano, Sedigheh Eslami, Daniel
Hodczak, Liu Liu, Andres Sevtsuk, and Gerard de Melo.
arXiv preprint arXiv:2406.01551.
2023
Inferior alveolar canal automatic detection with deep learning CNNs on CBCTs:
development of a novel model and release of open-source dataset and algorithm
Mattia Di Bartolomeo, Arrigo Pellacani, Federico Bolelli,
Marco Cipriano, Luca Lumetti, Sara Negrello, Stefano Allegretti,
Paolo Minafra, Federico Pollastri, Riccardo Nocini, and others.
Applied Sciences, 13(5), 3271.
2022
Improving segmentation of the inferior alveolar nerve through deep label
propagation
Marco Cipriano, Stefano Allegretti, Federico Bolelli, Federico
Pollastri, and Costantino Grana.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 21137-21146.
2022
Long-range 3D self-attention for MRI prostate segmentation
Federico Pollastri, Marco Cipriano, Federico Bolelli, and
Costantino Grana.
2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), 1-5.
2022
Deep segmentation of the mandibular canal: a new 3D annotated dataset of CBCT
volumes
Marco Cipriano, Stefano Allegretti, Federico Bolelli, Mattia Di
Bartolomeo, Federico Pollastri, Arrigo Pellacani, Paolo Minafra, Alexandre Anesi,
and Costantino Grana.
IEEE Access, 10, 11500-11510.
2021
The color out of space: learning self-supervised representations for earth
observation imagery
Stefano Vincenzi, Angelo Porrello, Pietro Buzzega, Marco Cipriano,
Pietro Fronte, Roberto Cuccu, Carla Ippoliti, Annamaria Conte, and Simone
Calderara.
2020 25th International Conference on Pattern Recognition (ICPR), 3034-3041.
2021
A cone beam computed tomography annotation tool for automatic detection of the
inferior alveolar nerve canal
Cristian Mercadante, Marco Cipriano, Federico Bolelli, Federico
Pollastri, Alexandre Anesi, Costantino Grana, and others.
Proceedings of the 16th International Joint Conference on Computer Vision,
Imaging and Computer Graphics Theory and Applications, Volume 4: VISAPP, 724-731.