The researchers proposed a biologically inspired framework that uses motion boundaries to learn object-centered visual representations from raw videos for individual images. This method combines readily available optical flow with clustering to generate pseudo-instance masks, and进行监督 single-image encoders for pixel-level pairwise metric learning without the need for manual annotation or camera calibration. The team first extracted 195 million pseudo-labeled frames from 7,163 hours of driving and network videos, and then expanded the self-training using motion verification to 421 million frames. The trained encoders reached Swin-H accuracy level and performed representation distillation across a series of Swin backbone networks. In tasks such as single-eye depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the model performed better or comparable to supervised and self-supervised pre-trained baselines, especially showing significant transfer effects in geometric and instance-sensitive tasks.
The researchers proposed a biologically inspired framework that uses motion boundaries to learn object-centered visual representations from individual images in the original video. This method combines readily available optical flow and clustering to generate pseudo-instance masks, and进行监督 single-image encoding to perform pixel-level pairwise metric learning without the need for manual annotation or camera calibration. The team first extracted 195 million pseudo-labeled frames from 7,163 hours of driving and network videos, and then expanded the self-training with motion verification to 421 million frames. The trained encoder reached Swin-H accuracy and performed representation distillation across a series of Swin backbone networks. In tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the model performed better or comparable to supervised and self-supervised pre-trained baselines, especially showing significant transfer effects in geometric and instance-sensitive tasks.