PureByte: SLMs Are the Future

We hold that most automated decisions over data (is there a credential in this file, which bytes of this ticket are personal data, what is this blob) are best made by small models trained for each job, and we present PureByte, an architecture, runtime and training stack for them. We use “small language model” in a broad sense: an AI Specialist is a small learned model of byte sequences that answers with decisions, scores, labels and byte ranges instead of generated text. It is a stack of ternary state-space blocks of the Mamba-2 family over a byte embedding and optional hashed n-gram tables, with heads decoded under constraints, stored as one file of 1 to 29 MiB that a dependency-free C++ runtime executes on ordinary CPUs, with bit-identical results across the kernels and thread counts of a platform. Specialists are trained from scratch in minutes to an hour on one consumer GPU by a factory that blocks leaks between training data and examinations and verifies that the runtime reproduces the trained model (its leak checks came after the three released examples, whose overlaps with their examinations we measure); an AI coding agent can drive it from a description of the job, which we call Vibe Training. We formulate the approach (the answer types, an argument for when a specialist needs no general knowledge, and a cost model that includes the cost of errors) and measure three released examples, which all answer with byte ranges gated by a decision (the other answer types have no trained model yet). On 164 of the 168 repositories of CredData’s test half, a 1.9-million-parameter model finds 3.5 times as many labeled credential lines as gitleaks with its default rules, at a close precision on labeled lines (3.1 times on all 168), after learning partly from labeled lines of the benchmark’s other half; with the benchmark’s obfuscated values replaced by the original ones, 2.9 times. An 8.3-million-parameter model scores F2 0.769 on the PII Masking Benchmark, 13th of 36 and above OpenAI’s Privacy Filter (1.4 billion parameters, about 50 million active, with a narrower policy), after training on the sources of four of its six tasks (on the other two tasks the two models are level); the best small pretrained encoders score higher at a few times its cost per byte. A 64-byte decision takes 0.8 to 1.5 ms on eight threads of a desktop CPU, 3 to 6 ms on one core. Each released model missed a pre-registered criterion as first written, and we report every one. We also report research on an attachable exact document memory, withdrawn from one application after a real-data test. Code (runtime and training stack): Apache-2.0. Weights of the three examples: PureByte Model License 1.0. This paper: CC BY 4.0.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23020056
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

PureByte: SLMs Are the Future

Pablo Sirvent Jiménez
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

PureByte: SLMs Are the Future

Pablo Sirvent Jiménez
preprint en

Abstract

We hold that most automated decisions over data (is there a credential in this file, which bytes of this ticket are personal data, what is this blob) are best made by small models trained for each job, and we present PureByte, an architecture, runtime and training stack for them. We use “small language model” in a broad sense: an AI Specialist is a small learned model of byte sequences that answers with decisions, scores, labels and byte ranges instead of generated text. It is a stack of ternary state-space blocks of the Mamba-2 family over a byte embedding and optional hashed n-gram tables, with heads decoded under constraints, stored as one file of 1 to 29 MiB that a dependency-free C++ runtime executes on ordinary CPUs, with bit-identical results across the kernels and thread counts of a platform. Specialists are trained from scratch in minutes to an hour on one consumer GPU by a factory that blocks leaks between training data and examinations and verifies that the runtime reproduces the trained model (its leak checks came after the three released examples, whose overlaps with their examinations we measure); an AI coding agent can drive it from a description of the job, which we call Vibe Training. We formulate the approach (the answer types, an argument for when a specialist needs no general knowledge, and a cost model that includes the cost of errors) and measure three released examples, which all answer with byte ranges gated by a decision (the other answer types have no trained model yet). On 164 of the 168 repositories of CredData’s test half, a 1.9-million-parameter model finds 3.5 times as many labeled credential lines as gitleaks with its default rules, at a close precision on labeled lines (3.1 times on all 168), after learning partly from labeled lines of the benchmark’s other half; with the benchmark’s obfuscated values replaced by the original ones, 2.9 times. An 8.3-million-parameter model scores F2 0.769 on the PII Masking Benchmark, 13th of 36 and above OpenAI’s Privacy Filter (1.4 billion parameters, about 50 million active, with a narrower policy), after training on the sources of four of its six tasks (on the other two tasks the two models are level); the best small pretrained encoders score higher at a few times its cost per byte. A 64-byte decision takes 0.8 to 1.5 ms on eight threads of a desktop CPU, 3 to 6 ms on one core. Each released model missed a pre-registered criterion as first written, and we report every one. We also report research on an attachable exact document memory, withdrawn from one application after a real-data test. Code (runtime and training stack): Apache-2.0. Weights of the three examples: PureByte Model License 1.0. This paper: CC BY 4.0.

Zenodo (CERN European Organization for Nuclear Research)
Incyte (United States) (US)
Peace, Justice and strong institutions
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.