Injective attention neural operators for partial differential equations on large-scale problems
Partial differential equations (PDEs) play a central role in modeling physical phenomena across science and engineering, including fluid dynamics, solid mechanics, electromagnetics, and multiphysics systems. Repeatedly solving PDEs over varying inputs, geometries, and parameters remains a fundamental yet computationally demanding task in large-scale simulations. Recently, neural operators have introduced a new paradigm for learning solution operators of PDEs, enabling mappings between infinite-dimensional function spaces. However existing transformer-based designs suffer from an intrinsic trade-off between scalability and accuracy. Quadratic-cost softmax attention achieves high accuracy but becomes computationally prohibitive on large-scale meshes, whereas linear attention variants reduce complexity to scale linearly with sequence length, but often at the cost of expressive power. In this paper, to resolve the above-mentioned limitations, we propose a new class of neural operators, termed Injective Attention Neural Operator (IANO), which reformulates attention as an injective mapping between functional representations relevant to PDE benchmarks. This injective formulation enables IANO to retain the global expressivity of softmax attention while achieving linear computational complexity. Theoretically, we rigorously prove a universal approximation theorem for IANO, demonstrating its ability to approximate arbitrary continuous nonlinear operators between Sobolev spaces. Empirically, IANO achieves state-of-the-art performance on seven out of eight standard PDE benchmarks, surpassing existing methods by a significant margin on those cases while performing comparably on the remaining one. Moreover, IANO generalizes effectively to complex, Reynolds-averaged Navier–Stokes-governed mesh datasets (e.g., car and airfoil designs), delivering superior accuracy and robustness beyond classical tasks. Through the synergy of theoretical guarantees and refined architectural design, IANO establishes a new foundation for accurate, efficient, and mathematically principled PDE operator learning within scientific machine learning.
Authors
- Weifang Weng (ORCID: https://orcid.org/0009-0002-1747-6286)
- Zhenya Yan (ORCID: https://orcid.org/0000-0002-9475-3753)
- Ming Zhong (ORCID: https://orcid.org/0000-0002-1671-920X)
- Lu Lu
Institutions
- Zhongyuan University of Technology (CN)
- University of Electronic Science and Technology of China (CN)
- Chinese Academy of Sciences (CN)
- Yale University (US)
- Academy of Mathematics and Systems Science (CN)
- University of Chinese Academy of Sciences (CN)
Publication Details
- Journal
- Journal of Computational Physics
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1016/j.jcp.2026.115457
- Primary Topic
- Model Reduction and Neural Networks
- Type
- article
- Field-Weighted Citation Impact
- 0.00