First, please prepare the image data following this instruction in LISA. We introduce the video datasets used in this project. Note that the data paths for video datasets are currently hard-coded in ...
Abstract: Surveillance videos are important for public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to ...
SlowFast-LLaVA is a training-free multimodal large language model (LLM) for video understanding and reasoning. Without requiring fine-tuning on any data, it achieves comparable or even better ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results