VITRA is a novel approach for pretraining Vision-Language-Action (VLA) models for robotic manipulation using large-scale, unscripted, real-world videos of human hand activities. Treating human hand as ...
First, please prepare the image data following this instruction in LISA. We introduce the video datasets used in this project. Note that the data paths for video datasets are currently hard-coded in ...