FAMoE-ST: Hierarchically Frozen Attention with Hybrid-Memory Experts for Traffic-Flow Forecasting
Accurate multi-step traffic-flow forecasting requires dynamic spatial modeling and adaptation to heterogeneous temporal and node-level patterns. This study proposes FAMoE-ST, a hierarchically frozen attention network with hybrid-memory experts. Historical flow, temporal context, and node identity are encoded as sensor tokens and processed by a six-layer Transformer initialized from GPT-2. To avoid dependence on arbitrary sensor indexing, the causal mask is replaced by all-to-all bidirectional spatial attention and every sensor uses the same GPT position ID. The lower four blocks are frozen, while the upper two blocks are adapted and their feed-forward networks are replaced by four-expert modules. A linear router is fused with a 32-slot memory router; each token retrieves four slots and activates two experts. With 12 observations predicting the next 12 steps, FAMoE-ST achieves MAE/RMSE/MAPE of 18.55/30.35/12.91% on PEMS04 and 14.57/24.08/9.71% on PEMS08. Relative to ST-LLM, MAE decreases by 6.97% and 7.39%, respectively. Component and capacity-matched controls support the roles of restricted adaptation, sparse expert capacity, and memory routing. A separate PEMS04 initialization control finds no advantage from GPT-2 pretraining over random initialization; the contribution is therefore attributed to the proposed structural adaptation rather than to transferred linguistic knowledge.
Authors
- Chenlong Li (ORCID: https://orcid.org/0000-0001-6947-4093)
- Fenghua Zhu (ORCID: https://orcid.org/0000-0003-2886-6968)
- Zhixue Wang (ORCID: https://orcid.org/0009-0009-6296-6146)
Institutions
- Chinese Academy of Sciences (CN)
- Shandong Jiaotong University (CN)
- Institute of Automation (CN)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/app16189140
- Primary Topic
- Traffic Prediction and Management Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00