Shortcut-Aware Logic-Sinking Distillation with Budget-Adaptive Decision Control for Efficient Long-Horizon Reasoning
Long-horizon reasoning improves complex question answering but often requires expensive chain-of-thought generation and may transfer shortcut-sensitive fragments during distillation. This paper addresses the focused problem of learning a compact student model that preserves decision-relevant logic while reducing shortcut reliance and inference cost. We propose SALS-FD, whose principal contribution is a reliability-weighted logic-sinking distillation objective: teacher traces are represented as logical units, each unit is evaluated for causal contribution and shortcut risk, and the reliable information is compressed into a latent student target. A budget-adaptive four-action controller is included as a complementary inference mechanism that decides whether to continue, verify, decide, or use a longer student-side reasoning path. On GSM8K, MATH, StrategyQA, and HotpotQA, the reported three-seed means reach an average Accuracy of 72.29%, Reasoning-Step F1 of 69.14%, and ECE of 7.15%, together with lower reasoning-token use and latency than the evaluated baselines. The simultaneous changes are interpreted as coupled consequences of filtering the shared latent reasoning target and calibrating the inference policy; they are empirical observations on the reported benchmarks rather than a general guarantee of trade-off-free improvement.
Authors
- Jiangbo Feng (ORCID: https://orcid.org/0009-0009-0226-8557)
- Qi Zeng
- Jiayang Lai
- Shifeng Lin
- Yuxiang He
- Jinchao Guo
- Ke Wang
Institutions
- Twitter (United States) (US)
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1142/s0218001426510183
- Primary Topic
- Intelligent Tutoring Systems and Adaptive Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00