Mathematically-Guided Detection of Floating-Point Errors
Floating-point computations are important for modern scientific and engineering software, especially for safety-critical systems, yet only a small subset of inputs typically trigger substantial numerical errors. Detecting such error-inducing inputs and the underlying bugs is therefore essential for improving their security and reliability. Existing techniques commonly rely on either oracle-driven exploration that repeatedly compares against high-precision references or search-driven heuristics. Despite the improvements made, they remain limited by (1) Expensive computation of high-precision oracles and (2) Lack of long-range convergence , which often requires dense probing near narrow error-inducing regions and expensive computation. We propose MGDE ( M athematically- G uided D etection of floating-point E rrors), a method that replaces trial-and-error exploration with mathematically defined targets and directed convergence. MGDE first uses condition-number theory to identify numerically unstable atomic operations without invoking expensive high-precision oracles during exploration. MGDE exploits the observation that extreme condition numbers occur near structured boundaries (e.g., cancellation points and singularities), reformulating detection as a numerical root-finding problem. By solving the resulting objectives with the Newton–Raphson method, MGDE can steer inputs toward error-prone regions from far-away initializations. We evaluate MGDE on GNU Scientific Library (GSL) functions and compare against two state-of-the-art baselines, ATOMU and FPCC, using triggered bugs as the primary metric. On 88 single-input functions, MGDE triggers 80 numerically validated bugs across 47 functions, outperforming ATOMU (70 bugs in 46 functions) and FPCC (53 bugs in 42 functions). MGDE is also faster: ATOMU and FPCC require 42.71× and 11.17× the exploration time of MGDE, respectively. Regarding multi-input functions, we evaluate MGDE under two complementary settings. On the native multi-input dataset of FPCC, MGDE detects 28 triggered bugs, while FPCC finds 23 bugs. MGDE also takes 8.91 seconds in total, compared with 2,100 seconds used by FPCC. On an additional external benchmark of 18 dual-input GSL functions, MGDE detects nine bugs not found by FPCC. Overall, MGDE substantially advances the state-of-the-art in both effectiveness and efficiency, and we report 16 previously unknown GSL bugs, which have been confirmed by the GSL community.
Authors
- Zhanwei Zhang (ORCID: https://orcid.org/0009-0002-5321-1787)
- Youshuai Tan (ORCID: https://orcid.org/0000-0002-3390-4079)
- Weiyi Shang (ORCID: https://orcid.org/0000-0001-6222-7444)
- Haonan Zhang (ORCID: https://orcid.org/0000-0002-6874-5581)
- Zishuo Ding (ORCID: https://orcid.org/0000-0002-0803-5609)
- Jinfu Chen (ORCID: https://orcid.org/0000-0001-7410-9146)
- Lianyu Zheng (ORCID: https://orcid.org/0009-0008-5872-1616)
Institutions
- University of Waterloo (CA)
- Wuhan University (CN)
- The Hong Kong University of Science and Technology (Guangzhou) (CN)
Publication Details
- Journal
- Proceedings of the ACM on software engineering.
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1145/3832291
- Primary Topic
- Numerical Methods and Algorithms
- Type
- article
- Field-Weighted Citation Impact
- 0.00