Enhancing transformer language models with liquid time-constant networks
The feedforward sublayer in a transformer block applies an identical, position-wise transformation to every token, with no mechanism to vary its computational behaviour based on syntactic role or local context. This paper investigates whether replacing these static feedforward layers with Closed-form Continuous-time (CfC) liquid layers is feasible within an autoregressive language model and whether the resulting model learns interpretable dynamics. Using knowledge distillation from GPT-2 followed by supervised fine-tuning, a 28.81M-parameter hybrid model has been trained on WikiText-103. The model achieves a word-level perplexity of 3262.47 versus 3585.52 for an identically parameterised MLP baseline, and scores of 26.14%, 52.18%, 26.68%, and 24.65% on HellaSwag, PIQA, ARC-Easy, and MMLU respectively. More importantly, inference-time probing of the model's single learned interpolation gate α ( x ) — the quantity the CfC cell actually computes in its default (non-ODE) operating mode — reveals a statistically significant layer-dependent effect (one-way ANOVA, F = 78.6 , p < 10 − 40 , η 2 = 0.349 ); a complementary test found no dependence on token level syntactic complexity ( F = 0.024 , p = 0.976 ). These findings indicate that the gate encodes some layer specific structure, but at a scale and with a pattern more modest than a smooth, monotonic timescale hierarchy. Training remains stable throughout, with gradient norms settling to 0.5–0.6 during fine-tuning.
Authors
- J. Anitha (ORCID: https://orcid.org/0000-0002-0293-0821)
- Nidhin Paul (ORCID: https://orcid.org/0000-0003-3602-6256)
- Shanu Mathew
Institutions
- Karunya University (IN)
Publication Details
- Journal
- Journal of Intelligent & Fuzzy Systems
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1177/18758967261494074
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00