Retrofitting language models to operate over bytes

Abstract Recent advances in artificial intelligence (AI) have largely been driven by large language models, deep neural networks that operate over discrete units called tokens. To represent text, most large language models use words or word fragments as the tokens, known as subword tokenization 1 . Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences—where meaning depends on the individual characters or bytes 2 . Models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models in performance. Here we introduce a general method for creating byte-level large language models through byteification that approach the capabilities of subword-based systems. We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training. The resulting models outperform earlier byte-level approaches and excel on character-level reasoning tasks, achieving practical inference speeds by efficiently processing byte-level information and adaptability by reusing the existing ecosystem around the source large language model. Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding.

Authors

Institutions

Publication Details

Journal
Nature
Published
2026-10-07
DOI
https://doi.org/10.1038/s41586-026-11111-4
Citations
1
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
3.77
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Retrofitting language models to operate over bytes

Valentin Hofmann, Tomasz Limisiewicz, Tyler Murray, Luca Soldaini et al.
1 citations
Nature
Topic Modeling
3.77
article

Retrofitting language models to operate over bytes

Valentin Hofmann, Tomasz Limisiewicz, Tyler Murray, Luca Soldaini, E De Ponti, Luke Zettlemoyer, Benjamin Minixhofer, Anna Korhonen, Noah A. Smith
article en
1 citations

Abstract

Abstract Recent advances in artificial intelligence (AI) have largely been driven by large language models, deep neural networks that operate over discrete units called tokens. To represent text, most large language models use words or word fragments as the tokens, known as subword tokenization 1 . Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences—where meaning depends on the individual characters or bytes 2 . Models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models in performance. Here we introduce a general method for creating byte-level large language models through byteification that approach the capabilities of subword-based systems. We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training. The resulting models outperform earlier byte-level approaches and excel on character-level reasoning tasks, achieving practical inference speeds by efficiently processing byte-level information and adaptability by reusing the existing ecosystem around the source large language model. Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding.

Nature
University of Washington (US), University of Cambridge (GB), Allen Institute for Artificial Intelligence (US), Munich Center for Machine Learning, Imperial College London (GB), Ludwig-Maximilians-Universität München (DE)
Openalex Percentile: Top 5%
Topic Modeling
3.77
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.