To address the issue that visual imitation strategies fail when encountering visually similar objects or containers, researchers used motion segmentation and the Transformer (ACT) system to introduce distractors that controlled the similarity of colors and shapes, accurately pinpointing the failure mode to the grasping and placement stages. The study found that the sensitivity to distractions is specific to the type of visual similarity and the operational stage. To address this, the team evaluated complementary interventions such as distraction enhancement, stage-dependent attention regularization, and appearance-based visual prompts, which significantly improved the robustness of the strategy in both simulation environments and the physical UR3e robot. Additionally, the study confirmed that this failure mode also occurs in state-conditioning tasks involving pre-trained visual-language-motion strategies (such as medical instrument operation). The results indicate that even when the underlying operational skills are intact, visual distractors can still lead to incorrect object or destination selection; explicitly improving target selection can significantly restore performance under different visual-motor strategy learning paradigms.