To overcome limitations such as errors in diffusion model texts and lack of global aesthetics in code generation, the study proposes a new paradigm of “editable visual design” driven by Coding Agents. This approach uses Visual Language Models (VLM) to serve as the creative brain for demand understanding and aesthetic judgment, utilizing image generation models to build visual world simulators for synthesizing independent assets. The system employs a “think before acting” closed-loop workflow: Agents generate isolated assets, write native HTML/CSS, and iterate and optimize based on rendering feedback, ultimately delivering editable outputs with decoupled layers and real text. Users can intuitively perform mouse dragging and layout adjustments in the graphical interface. In verification scenarios such as posters and infographics, this paradigm successfully achieves both fine aesthetics and production-level editability.