X-Ternary: A Deterministic Causal 2-Bit Inference Engine for Large Language Models
Current Large Language Models (LLMs) rely on associative matrix multiplication and probabilistic token generation, which inherently leads to logical hallucinations and computational inefficiencies, particularly in deterministic fields such as mathematics and code generation. We present X-Ternary, a novel hybrid inference and training architecture that radically compresses LLM weights to 2-bit ternary states {-1, 0, 1} while enforcing strict causal non-associativity directly at the silicon level on NVIDIA GPUs. By utilizing 2:4 Structured Sparsity, block-level FP8 scaling for outlier preservation, and dynamic hardware routing between Tensor Cores and CUDA ALU Cores, X-Ternary achieves exact neuro-symbolic logic. By grounding the computation in intensional mathematics—where the computational process and causal history (identity of origin) are inseparable from the final state—X-Ternary eliminates the associative probability flaws of traditional GEMM operations. Furthermore, we introduce a dual-phase Curriculum Learning pipeline (Brain Stem Locking) utilizing Reinforcement Learning (RL) with Straight-Through Estimators (STE) and Elastic Weight Consolidation (EWC) to freeze foundational logical weights. The finalized architecture reduces memory constraints by approximately 80%, enabling 70B-parameter class models to run on single consumer-grade GPUs with near-zero latency mathematical precision.
Authors
- Michal Mazgal
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22877344
- Primary Topic
- Machine Learning in Materials Science
- Type
- preprint