RefDiT: Local Attribute Guidance in Reference-Based Image Generation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
RefDiT is a new framework based on reference images, designed to address the issue of difficulty in accurately locating related elements in multi-object scenarios using existing methods. This method uses local region attributes to construct perceptual condition signals by inputting reference images, text prompts, and optional user-guided context. In implementation, RefDiT decomposes identifier tokens at the attribute level to generate condition signals and adjusts the context during inference prompts, which is used to train the LoRA blocks of the generative model based on Diffusion Transformer (DiT). By establishing a correspondence between identifier tokens and local regions of reference images, RefDiT achieves more effective local guidance control.