AuraTracer智迹闻
中文

EVENT DOSSIER

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
6mentions
SummaryAI generated

The researchers proposed a unified hierarchical visual Transformer (ViT) framework called BMT, aimed at simultaneously addressing the tasks of translating SAR images into optical images and performing semantic segmentation. This framework optimizes both tasks through a shared hierarchical visual Transformer, integrating local ViT blocks, an enhanced output module, a control network-based conditional injection mechanism, and a bounded Kendall uncertainty weighting scheme. Experimental evaluations were conducted on the WHU-OPT-SAR paired dataset and the antidual datasets built based on HRSID and DIOR, and the results show that this method is competitive in terms of image translation quality and segmentation performance. Currently, the relevant datasets and source code are available publicly.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Bridging Modalities and TasksControlNetDIORHRSIDLocalViTBlockWHU-OPT-SAR

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Bridging Modalities and…1Bridging Modalities and…1Bridging Modalities and…1Bridging Modalities and…1Bridging Modalities and…1ControlNet × DIOR1

SignalsSIGNALS

Keyword heat
  • Bridging Modalities and Tasks1
  • LocalViTBlock1
  • ControlNet1
  • WHU-OPT-SAR1
  • HRSID1
  • DIOR1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

提出一种名为 BMT 的统一分层 ViT 框架,通过共享层级视觉 Transformer 联合优化 SAR 到光学图像翻译与语义分割任务。该框架集成局部 ViT 块、增强输出模块、控制网式条件注入机制及有界肯德尔不确定性加权方案,在 WHU-OPT-SAR 配对数据集及基于 HRSID 和 DIOR 构建的无对偶船只数据集上评估。实验表明该方法在翻译质量与分割性能上表现具有竞争力,相关数据集与源代码已公开。