The researchers proposed Repeat-After-Me, a black-box adaptive visual prompt injection attack that requires no semantic correlation with the target model or oral authorization. Experiments showed that the attack’s success rate on visual language models such as Qwen3.6-27B and GPT-5.5 exceeded 80% and 47%, respectively. Even after optimizing the single agent model, the attack still maintained a 43%-46% original success rate for both commercial victims, with cross-sample transferability remaining at 64%-66%. The researchers verified the effectiveness of the attack in the OpenClaw real environment, indicating that untrusted users could enable sensitive behaviors such as remote code execution and secret leakage by minimizing the size of the injected images to cover TOOLS.md. This attack vector remains effective even when adaptive text prompt injection fails.