Enhancing transformer language models with liquid time-constant networks

The feedforward sublayer in a transformer block applies an identical, position-wise transformation to every token, with no mechanism to vary its computational behaviour based on syntactic role or local context. This paper investigates whether replacing these static feedforward layers with Closed-form Continuous-time (CfC) liquid layers is feasible within an autoregressive language model and whether the resulting model learns interpretable dynamics. Using knowledge distillation from GPT-2 followed by supervised fine-tuning, a 28.81M-parameter hybrid model has been trained on WikiText-103. The model achieves a word-level perplexity of 3262.47 versus 3585.52 for an identically parameterised MLP baseline, and scores of 26.14%, 52.18%, 26.68%, and 24.65% on HellaSwag, PIQA, ARC-Easy, and MMLU respectively. More importantly, inference-time probing of the model's single learned interpolation gate α ( x ) — the quantity the CfC cell actually computes in its default (non-ODE) operating mode — reveals a statistically significant layer-dependent effect (one-way ANOVA, F = 78.6 , p < 10 − 40 , η 2 = 0.349 ); a complementary test found no dependence on token level syntactic complexity ( F = 0.024 , p = 0.976 ). These findings indicate that the gate encodes some layer specific structure, but at a scale and with a pattern more modest than a smooth, monotonic timescale hierarchy. Training remains stable throughout, with gradient norms settling to 0.5–0.6 during fine-tuning.

Authors

Institutions

Publication Details

Journal
Journal of Intelligent & Fuzzy Systems
Published
2026-10-08
DOI
https://doi.org/10.1177/18758967261494074
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Enhancing transformer language models with liquid time-constant networks

J. Anitha, Nidhin Paul, Shanu Mathew
Journal of Intelligent & Fuzzy Systems
Topic Modeling
article

Enhancing transformer language models with liquid time-constant networks

J. Anitha, Nidhin Paul, Shanu Mathew
article en

Abstract

The feedforward sublayer in a transformer block applies an identical, position-wise transformation to every token, with no mechanism to vary its computational behaviour based on syntactic role or local context. This paper investigates whether replacing these static feedforward layers with Closed-form Continuous-time (CfC) liquid layers is feasible within an autoregressive language model and whether the resulting model learns interpretable dynamics. Using knowledge distillation from GPT-2 followed by supervised fine-tuning, a 28.81M-parameter hybrid model has been trained on WikiText-103. The model achieves a word-level perplexity of 3262.47 versus 3585.52 for an identically parameterised MLP baseline, and scores of 26.14%, 52.18%, 26.68%, and 24.65% on HellaSwag, PIQA, ARC-Easy, and MMLU respectively. More importantly, inference-time probing of the model's single learned interpolation gate α ( x ) — the quantity the CfC cell actually computes in its default (non-ODE) operating mode — reveals a statistically significant layer-dependent effect (one-way ANOVA, F = 78.6 , p < 10 − 40 , η 2 = 0.349 ); a complementary test found no dependence on token level syntactic complexity ( F = 0.024 , p = 0.976 ). These findings indicate that the gate encodes some layer specific structure, but at a scale and with a pattern more modest than a smooth, monotonic timescale hierarchy. Training remains stable throughout, with gradient norms settling to 0.5–0.6 during fine-tuning.

Journal of Intelligent & Fuzzy Systems
Karunya University (IN)
Openalex Percentile: Top 12%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.