Research

About

I am a PhD student at the Hasso Plattner Institute (HPI), working with Prof. Gerard de Melo. My research lies at the intersection of vision-language models, computer vision, and vector graphics. I am interested in studying where current VLMs fall short in visually and relationally complex settings, and in teaching models to generate and reason with visual content in ways that more closely resemble how humans think and communicate.

I am currently working with Prof. Yael Vinker and Prof. Antonio Torralba (MIT CSAIL) on a project to enable sketching capabilities in the model's reasoning process, with the goal of building systems that can explain concepts visually and collaborate with students in a shared whiteboard space. I also worked with Prof. Andres Sevtšuk (MIT DUSP) on a project to detect social interactions in semantically complex urban scenes, with the goal of mapping which areas of modern cities are more socially vibrant than others.

Latest News

  1. I am spending three months at MIT CSAIL as a visiting researcher in Prof. Antonio Torralba’s lab.
  2. Our work MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes was accepted at AAAI 2026.
  3. Our project Whiteboard-literate AI won a grant from the HPI–MIT Designing for Sustainability program.
  4. Our work Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision was accepted at ICML 2025.

Publications

2026

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

Liu Liu, Alexandra Schild, Marco Cipriano, Fatimeh Al Ghannam, Freya Tan, Gerard de Melo, and Andres Sevtsuk.

Proceedings of the AAAI Conference on Artificial Intelligence, 40(45), 38935-38942.

2025

Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision

Marco Cipriano, Moritz Feuerpfeil, and Gerard de Melo.

Forty-second International Conference on Machine Learning.

2024

ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection

Maryam Hosseini, Marco Cipriano, Sedigheh Eslami, Daniel Hodczak, Liu Liu, Andres Sevtsuk, and Gerard de Melo.

arXiv preprint arXiv:2406.01551.

2023

Inferior alveolar canal automatic detection with deep learning CNNs on CBCTs: development of a novel model and release of open-source dataset and algorithm

Mattia Di Bartolomeo, Arrigo Pellacani, Federico Bolelli, Marco Cipriano, Luca Lumetti, Sara Negrello, Stefano Allegretti, Paolo Minafra, Federico Pollastri, Riccardo Nocini, and others.

Applied Sciences, 13(5), 3271.

2022

Improving segmentation of the inferior alveolar nerve through deep label propagation

Marco Cipriano, Stefano Allegretti, Federico Bolelli, Federico Pollastri, and Costantino Grana.

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21137-21146.

2022

Long-range 3D self-attention for MRI prostate segmentation

Federico Pollastri, Marco Cipriano, Federico Bolelli, and Costantino Grana.

2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), 1-5.

2022

Deep segmentation of the mandibular canal: a new 3D annotated dataset of CBCT volumes

Marco Cipriano, Stefano Allegretti, Federico Bolelli, Mattia Di Bartolomeo, Federico Pollastri, Arrigo Pellacani, Paolo Minafra, Alexandre Anesi, and Costantino Grana.

IEEE Access, 10, 11500-11510.

2021

The color out of space: learning self-supervised representations for earth observation imagery

Stefano Vincenzi, Angelo Porrello, Pietro Buzzega, Marco Cipriano, Pietro Fronte, Roberto Cuccu, Carla Ippoliti, Annamaria Conte, and Simone Calderara.

2020 25th International Conference on Pattern Recognition (ICPR), 3034-3041.

2021

A cone beam computed tomography annotation tool for automatic detection of the inferior alveolar nerve canal

Cristian Mercadante, Marco Cipriano, Federico Bolelli, Federico Pollastri, Alexandre Anesi, Costantino Grana, and others.

Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Volume 4: VISAPP, 724-731.