Activation function saturation determines robustness to conductance nonlinearity in analog in-memory learning

Abstract In analog memristive neural networks the weighted summation is performed efficiently in the crossbar, but how the activation function interacts with the imperfect, non-linear conductance updates of the devices remains poorly understood. We compare five activation functions—Sigmoid, Tanh, ReLU, GELU and SiLU—on MNIST and Fashion-MNIST, using an idealized software configuration and a hardware configuration parameterized by the measured conductance characteristics of a fabricated TiO 2 device. Learning rates are tuned separately for every function and configuration, and every result is averaged over three random seeds. Under this protocol all five functions are statistically indistinguishable in software, spanning 0.68 points on MNIST and 0.56 on Fashion-MNIST, so any separation on hardware cannot be attributed to one function being better suited to the task. On hardware they separate sharply: the software-to-hardware gap ranges from 0.23 ± 0.12 points for Tanh to 10.54 ± 1.23 for ReLU. The dividing line is saturation rather than the shape of the negative branch. GELU and SiLU, whose small attenuated negative response might be expected to mitigate the problem, degrade as severely as ReLU; the two saturating functions lose 1.36 points on average against 7.82 for the three non-saturating ones. The mechanism is measured directly: the cumulative weight-update magnitude spans almost three orders of magnitude across the five functions and predicts the gap with r² = 0.85. Saturating functions bound the update and keep the device within the near-linear region of its conductance response, while non-saturating functions drive updates that traverse the strongly nonlinear and asymmetric region. This hardware ranking inverts the one familiar from digital deep learning, and does so for the functions that current practice favors.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-25
DOI
https://doi.org/10.1038/s41598-026-72039-3
Primary Topic
Advanced Memory and Neural Computing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Activation function saturation determines robustness to conductance nonlinearity in analog in-memory learning

Tolga Aydın, Fatih Gül, Baki Gökgöz
Scientific Reports
Advanced Memory and Neural Computing
article

Activation function saturation determines robustness to conductance nonlinearity in analog in-memory learning

Tolga Aydın, Fatih Gül, Baki Gökgöz
article en

Abstract

Abstract In analog memristive neural networks the weighted summation is performed efficiently in the crossbar, but how the activation function interacts with the imperfect, non-linear conductance updates of the devices remains poorly understood. We compare five activation functions—Sigmoid, Tanh, ReLU, GELU and SiLU—on MNIST and Fashion-MNIST, using an idealized software configuration and a hardware configuration parameterized by the measured conductance characteristics of a fabricated TiO 2 device. Learning rates are tuned separately for every function and configuration, and every result is averaged over three random seeds. Under this protocol all five functions are statistically indistinguishable in software, spanning 0.68 points on MNIST and 0.56 on Fashion-MNIST, so any separation on hardware cannot be attributed to one function being better suited to the task. On hardware they separate sharply: the software-to-hardware gap ranges from 0.23 ± 0.12 points for Tanh to 10.54 ± 1.23 for ReLU. The dividing line is saturation rather than the shape of the negative branch. GELU and SiLU, whose small attenuated negative response might be expected to mitigate the problem, degrade as severely as ReLU; the two saturating functions lose 1.36 points on average against 7.82 for the three non-saturating ones. The mechanism is measured directly: the cumulative weight-update magnitude spans almost three orders of magnitude across the five functions and predicts the gap with r² = 0.85. Saturating functions bound the update and keep the device within the near-linear region of its conductance response, while non-saturating functions drive updates that traverse the strongly nonlinear and asymmetric region. This hardware ranking inverts the one familiar from digital deep learning, and does so for the functions that current practice favors.

Scientific Reports
Gümüşhane University (TR), Recep Tayyip Erdoğan University (TR), Atatürk University (TR)
Openalex Percentile: Top 21%
Advanced Memory and Neural Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.