AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence with a Native NVMe Hardware Memory Bus
AviGPT-250M-Instruct is a 250M-parameter autoregressive small language model (SLM) introducing a Semi-Parametric Decoupling paradigm for resource-constrained edge computing. Rather than overloading transformer weights with static encyclopedic memorization and floating-point arithmetic approximation, AviGPT-250M delegates factual storage to a native, sub-millisecond (0.0015 ms / 1.52 µs) NVMe Hardware Memory Bus powered by SQLite Full-Text Search (FTS5) with Okapi BM25 ranking, and delegates arithmetic evaluation to a sandboxed deterministic AST SafeMath evaluator. Across an identical 8-model competitive benchmark on an NVIDIA Tesla T4 GPU against models ranging from 125M to 1.1B parameters, AviGPT-250M achieves 100.0% Factual Accuracy and 100.0% Deterministic Math Precision, delivering the Global #1 Composite Efficiency score of 0.40 while fitting inside an ultra-compact resident VRAM footprint of just 488 MB.
Authors
- Avinash Ricky Yadlapalli
Institutions
- Oldham Council (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23067632
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint