To address the unreliability of existing multimodal agents in image generation and editing tasks due to the lack of verification from external knowledge, the WeAgent team proposed the WeAgent-MMGenEdit full-stack solution. This solution integrates the WeAgent-Harness multimodal runtime, a scalable data construction pipeline, and the bilingual benchmark WeBench-MMGenEdit, and employs a two-step training process based on supervised fine-tuning (SFT) and reinforcement learning (RL). Through this approach, researchers constructed a dataset containing 23,000 supervised trajectories and 14,700 reinforcement learning tasks, and trained an agent policy model with a total of 30 billion parameters and 3 billion active parameters. Experiments show that this model outperforms models of the same size in performance, approaching the level of agents with 100 billion parameters.