PureByte: SLMs Are the Future
We hold that most automated decisions over data (is there a credential in this file, which bytes of this ticket are personal data, what is this blob) are best made by small models trained for each job, and we present PureByte, an architecture, runtime and training stack for them. We use “small language model” in a broad sense: an AI Specialist is a small learned model of byte sequences that answers with decisions, scores, labels and byte ranges instead of generated text. It is a stack of ternary state-space blocks of the Mamba-2 family over a byte embedding and optional hashed n-gram tables, with heads decoded under constraints, stored as one file of 1 to 29 MiB that a dependency-free C++ runtime executes on ordinary CPUs, with bit-identical results across the kernels and thread counts of a platform. Specialists are trained from scratch in minutes to an hour on one consumer GPU by a factory that blocks leaks between training data and examinations and verifies that the runtime reproduces the trained model (its leak checks came after the three released examples, whose overlaps with their examinations we measure); an AI coding agent can drive it from a description of the job, which we call Vibe Training. We formulate the approach (the answer types, an argument for when a specialist needs no general knowledge, and a cost model that includes the cost of errors) and measure three released examples, which all answer with byte ranges gated by a decision (the other answer types have no trained model yet). On 164 of the 168 repositories of CredData’s test half, a 1.9-million-parameter model finds 3.5 times as many labeled credential lines as gitleaks with its default rules, at a close precision on labeled lines (3.1 times on all 168), after learning partly from labeled lines of the benchmark’s other half; with the benchmark’s obfuscated values replaced by the original ones, 2.9 times. An 8.3-million-parameter model scores F2 0.769 on the PII Masking Benchmark, 13th of 36 and above OpenAI’s Privacy Filter (1.4 billion parameters, about 50 million active, with a narrower policy), after training on the sources of four of its six tasks (on the other two tasks the two models are level); the best small pretrained encoders score higher at a few times its cost per byte. A 64-byte decision takes 0.8 to 1.5 ms on eight threads of a desktop CPU, 3 to 6 ms on one core. Each released model missed a pre-registered criterion as first written, and we report every one. We also report research on an attachable exact document memory, withdrawn from one application after a real-data test. Code (runtime and training stack): Apache-2.0. Weights of the three examples: PureByte Model License 1.0. This paper: CC BY 4.0.
Authors
- Pablo Sirvent Jiménez
Institutions
- Incyte (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23020055
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint