AuraTracer智迹闻
中文

EVENT DOSSIER

Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A controlled study on a variant of DeepONet revealed the significant impact of the attention mechanism on the accuracy of solving partial differential equations (PDEs). The researchers tested five different DeepONet variants with two architecture modes: data-driven and physical information-based, and evaluated them on benchmark problems such as source-driven transient one-dimensional nonlinear diffusion-reaction equations, transient one-dimensional viscous Burgers equations, and two-dimensional Poisson heat conduction problems. The results showed that models using a perceptron segmentation and cross-attention configuration reduced the average relative L_2 error by 2.4 to 28.0 times in all benchmark combinations, with the best configuration reducing it by 3.5 to 32.3 times; while the branch-of-attention mechanism using only dot-product fusion performed inconsistently. Although叠加 cross-attention improved all six scenarios, the improvement was minimal, global pre-mixing did not provide consistent benefits, and increasing the depth of cross-attention led to a significant increase in cost and a decrease in benefits. Overall, query-dependent cross-attention was proven to improve accuracy…

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures

A controlled study on a variant of DeepONet revealed the impact of the attention mechanism on the accuracy of PDE solutions. The researchers tested five DeepONet variants with different attention mechanisms in both data-driven and physical information modes, and evaluated them on benchmark problems such as source-driven transient one-dimensional nonlinear diffusion-reaction equations, transient one-dimensional viscous Burgers equations, and two-dimensional Poisson heat conduction problems. The results showed that configurations with sensor-based segmentation and cross-attention reduced the average relative L_2 error of the classical DeepONet by 2.4 to 28.0 times across all benchmark training combinations, with the best configuration reducing it by 3.5 to 32.3 times; whereas the branch-specific attention mechanism with only dot-product fusion performed inconsistently, and although adding cross-attention improved all six scenarios, the improvement was minimal. Global pre-mixing did not provide consistent benefits; increasing the depth of cross-attention could improve accuracy but at a significant cost, with diminishing returns. Overall, query-dependent cross-attention was the most…