Injective attention neural operators for partial differential equations on large-scale problems

Partial differential equations (PDEs) play a central role in modeling physical phenomena across science and engineering, including fluid dynamics, solid mechanics, electromagnetics, and multiphysics systems. Repeatedly solving PDEs over varying inputs, geometries, and parameters remains a fundamental yet computationally demanding task in large-scale simulations. Recently, neural operators have introduced a new paradigm for learning solution operators of PDEs, enabling mappings between infinite-dimensional function spaces. However existing transformer-based designs suffer from an intrinsic trade-off between scalability and accuracy. Quadratic-cost softmax attention achieves high accuracy but becomes computationally prohibitive on large-scale meshes, whereas linear attention variants reduce complexity to scale linearly with sequence length, but often at the cost of expressive power. In this paper, to resolve the above-mentioned limitations, we propose a new class of neural operators, termed Injective Attention Neural Operator (IANO), which reformulates attention as an injective mapping between functional representations relevant to PDE benchmarks. This injective formulation enables IANO to retain the global expressivity of softmax attention while achieving linear computational complexity. Theoretically, we rigorously prove a universal approximation theorem for IANO, demonstrating its ability to approximate arbitrary continuous nonlinear operators between Sobolev spaces. Empirically, IANO achieves state-of-the-art performance on seven out of eight standard PDE benchmarks, surpassing existing methods by a significant margin on those cases while performing comparably on the remaining one. Moreover, IANO generalizes effectively to complex, Reynolds-averaged Navier–Stokes-governed mesh datasets (e.g., car and airfoil designs), delivering superior accuracy and robustness beyond classical tasks. Through the synergy of theoretical guarantees and refined architectural design, IANO establishes a new foundation for accurate, efficient, and mathematically principled PDE operator learning within scientific machine learning.

Authors

Institutions

Publication Details

Journal
Journal of Computational Physics
Published
2026-10-06
DOI
https://doi.org/10.1016/j.jcp.2026.115457
Primary Topic
Model Reduction and Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Injective attention neural operators for partial differential equations on large-scale problems

Weifang Weng, Zhenya Yan, Ming Zhong, Lu Lu
Journal of Computational Physics
Model Reduction and Neural Networks
article

Injective attention neural operators for partial differential equations on large-scale problems

Weifang Weng, Zhenya Yan, Ming Zhong, Lu Lu
article en

Abstract

Partial differential equations (PDEs) play a central role in modeling physical phenomena across science and engineering, including fluid dynamics, solid mechanics, electromagnetics, and multiphysics systems. Repeatedly solving PDEs over varying inputs, geometries, and parameters remains a fundamental yet computationally demanding task in large-scale simulations. Recently, neural operators have introduced a new paradigm for learning solution operators of PDEs, enabling mappings between infinite-dimensional function spaces. However existing transformer-based designs suffer from an intrinsic trade-off between scalability and accuracy. Quadratic-cost softmax attention achieves high accuracy but becomes computationally prohibitive on large-scale meshes, whereas linear attention variants reduce complexity to scale linearly with sequence length, but often at the cost of expressive power. In this paper, to resolve the above-mentioned limitations, we propose a new class of neural operators, termed Injective Attention Neural Operator (IANO), which reformulates attention as an injective mapping between functional representations relevant to PDE benchmarks. This injective formulation enables IANO to retain the global expressivity of softmax attention while achieving linear computational complexity. Theoretically, we rigorously prove a universal approximation theorem for IANO, demonstrating its ability to approximate arbitrary continuous nonlinear operators between Sobolev spaces. Empirically, IANO achieves state-of-the-art performance on seven out of eight standard PDE benchmarks, surpassing existing methods by a significant margin on those cases while performing comparably on the remaining one. Moreover, IANO generalizes effectively to complex, Reynolds-averaged Navier–Stokes-governed mesh datasets (e.g., car and airfoil designs), delivering superior accuracy and robustness beyond classical tasks. Through the synergy of theoretical guarantees and refined architectural design, IANO establishes a new foundation for accurate, efficient, and mathematically principled PDE operator learning within scientific machine learning.

Journal of Computational PhysicsVol. 569
Zhongyuan University of Technology (CN), University of Electronic Science and Technology of China (CN), Chinese Academy of Sciences (CN), Yale University (US), Academy of Mathematics and Systems Science (CN), University of Chinese Academy of Sciences (CN)
Openalex Percentile: Top 10%
Model Reduction and Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.