MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
The research team proposed the MINT model, which directly generates hand trajectories in the world coordinate system based on self-view RGB videos for the first time. This model uses shared spatio-temporal video representations to jointly predict camera trajectories, in-frame hand states, and the presence of hands frame by frame, and outputs hand movements in the world space through explicit coordinate transformation. To address the shortage of annotations, the team developed the open-source annotation tool EGOPIPELINE, which converts public self-view videos into structured supervised data. MINT is first pre-trained on a large-scale pseudo-label dataset, and then fine-tuned on a small but high-quality joint annotation set. In public benchmark tests, the model demonstrated significant improvements in hand trajectory accuracy, camera trajectory estimation, and end-to-end generation speed, and could generalize to unseen self-view datasets with zero samples. Currently, the team has open-sourced the model code, inference code, annotation tool, and a curated self-view trajectory dataset containing 1,021 hours of data.