The Evolution of Large Language Model Architectures: From the Transformer to Hybrid Models (2017-2026)

Preprint, version 1.0 (September 2026) The Transformer of 2017 still sits at the core of every current large language model, yet almost everything around it has been rebuilt, some parts repeatedly. This survey traces the architectural development of text-only LLMs from 2017 to early 2026 as four epochs, each driven by a concrete bottleneck. In the exploration epoch (2018–2019), encoder-, decoder-, and encoder-decoder designs competed, and decoder-only prevailed for structural reasons that paid off only later. In the scaling epoch (2020–2022), capability was bought with parameters and data, while the first efficiency ideas failed for lack of a binding problem. In the efficiency epoch (2023–2024), inference cost became that problem: the field converged on the modern decoder recipe, attacked the KV cache factor by factor (GQA, MLA, sliding windows, paging), and revived mixture-of-experts and recurrent sequence mixing; nearly every mechanism was a rediscovery of a previously shelved idea. In the convergence epoch (2025–2026), independent model families arrived at the same design principle: a small minority of full-attention layers, interleaved with linear mixers and routed experts, retained to restore the exact recall that fixed-size states surrender. In parallel, reasoning models opened a test-time-compute axis whose cost lands on the same object: the memory a model carries per generated token. I synthesize four long-run development lines, compare the architecture families qualitatively, and state the open problems, including the 2026 dissent over the attention operator itself. Every core concept is explained once and in full, so the survey serves simultaneously as an introduction for students and a reference for researchers. Contents Introduction Foundations: The Transformer Epoch I: Exploration (2018–2019) Epoch II: Scaling (2020–2022) Epoch III: Efficiency and the Open Wave (2023–2024) Epoch IV: Convergence and New Axes (2025–2026) Synthesis: Nine Years in Four Lines Conclusion 48 pages, 14 figures, 9 tables, 116 references. Also available at https://finnrmn.com/papers/llm-architectures/.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22785090
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Evolution of Large Language Model Architectures: From the Transformer to Hybrid Models (2017-2026)

Finn Reimann
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

The Evolution of Large Language Model Architectures: From the Transformer to Hybrid Models (2017-2026)

Finn Reimann
preprint en

Abstract

Preprint, version 1.0 (September 2026) The Transformer of 2017 still sits at the core of every current large language model, yet almost everything around it has been rebuilt, some parts repeatedly. This survey traces the architectural development of text-only LLMs from 2017 to early 2026 as four epochs, each driven by a concrete bottleneck. In the exploration epoch (2018–2019), encoder-, decoder-, and encoder-decoder designs competed, and decoder-only prevailed for structural reasons that paid off only later. In the scaling epoch (2020–2022), capability was bought with parameters and data, while the first efficiency ideas failed for lack of a binding problem. In the efficiency epoch (2023–2024), inference cost became that problem: the field converged on the modern decoder recipe, attacked the KV cache factor by factor (GQA, MLA, sliding windows, paging), and revived mixture-of-experts and recurrent sequence mixing; nearly every mechanism was a rediscovery of a previously shelved idea. In the convergence epoch (2025–2026), independent model families arrived at the same design principle: a small minority of full-attention layers, interleaved with linear mixers and routed experts, retained to restore the exact recall that fixed-size states surrender. In parallel, reasoning models opened a test-time-compute axis whose cost lands on the same object: the memory a model carries per generated token. I synthesize four long-run development lines, compare the architecture families qualitatively, and state the open problems, including the 2026 dissent over the attention operator itself. Every core concept is explained once and in full, so the survey serves simultaneously as an introduction for students and a reference for researchers. Contents Introduction Foundations: The Transformer Epoch I: Exploration (2018–2019) Epoch II: Scaling (2020–2022) Epoch III: Efficiency and the Open Wave (2023–2024) Epoch IV: Convergence and New Axes (2025–2026) Synthesis: Nine Years in Four Lines Conclusion 48 pages, 14 figures, 9 tables, 116 references. Also available at https://finnrmn.com/papers/llm-architectures/.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The Evolution of Large Language Model Architectures: From the Transformer to Hybrid Models (2017-2026) — Finn Reimann · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS