3

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models

Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-time interactive environments impose new demands on VLMs, requiring them to generate utterances that …

SAB3R: Semantic-Augmented Backbone in 3D Reconstruction

We introduce a new task, Map and Locate, which unifies the traditionally distinct objectives of open-vocabulary segmentation - detecting and segmenting object instances based on natural language queries - and 3D reconstruction, the process of …