
VideoRefer Suite
·
DL·ML/Paper
https://arxiv.org/abs/2501.00599v1 VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMVideo Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, tharxiv.org arXiv에 241231에..