X-Ternary: A Deterministic Causal 2-Bit Inference Engine for Large Language Models

Current Large Language Models (LLMs) rely on associative matrix multiplication and probabilistic token generation, which inherently leads to logical hallucinations and computational inefficiencies, particularly in deterministic fields such as mathematics and code generation. We present X-Ternary, a novel hybrid inference and training architecture that radically compresses LLM weights to 2-bit ternary states {-1, 0, 1} while enforcing strict causal non-associativity directly at the silicon level on NVIDIA GPUs. By utilizing 2:4 Structured Sparsity, block-level FP8 scaling for outlier preservation, and dynamic hardware routing between Tensor Cores and CUDA ALU Cores, X-Ternary achieves exact neuro-symbolic logic. By grounding the computation in intensional mathematics—where the computational process and causal history (identity of origin) are inseparable from the final state—X-Ternary eliminates the associative probability flaws of traditional GEMM operations. Furthermore, we introduce a dual-phase Curriculum Learning pipeline (Brain Stem Locking) utilizing Reinforcement Learning (RL) with Straight-Through Estimators (STE) and Elastic Weight Consolidation (EWC) to freeze foundational logical weights. The finalized architecture reduces memory constraints by approximately 80%, enabling 70B-parameter class models to run on single consumer-grade GPUs with near-zero latency mathematical precision.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22877345
Primary Topic
Machine Learning in Materials Science
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

X-Ternary: A Deterministic Causal 2-Bit Inference Engine for Large Language Models

Michal Mazgal
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Materials Science
preprint

X-Ternary: A Deterministic Causal 2-Bit Inference Engine for Large Language Models

Michal Mazgal
preprint en

Abstract

Current Large Language Models (LLMs) rely on associative matrix multiplication and probabilistic token generation, which inherently leads to logical hallucinations and computational inefficiencies, particularly in deterministic fields such as mathematics and code generation. We present X-Ternary, a novel hybrid inference and training architecture that radically compresses LLM weights to 2-bit ternary states {-1, 0, 1} while enforcing strict causal non-associativity directly at the silicon level on NVIDIA GPUs. By utilizing 2:4 Structured Sparsity, block-level FP8 scaling for outlier preservation, and dynamic hardware routing between Tensor Cores and CUDA ALU Cores, X-Ternary achieves exact neuro-symbolic logic. By grounding the computation in intensional mathematics—where the computational process and causal history (identity of origin) are inseparable from the final state—X-Ternary eliminates the associative probability flaws of traditional GEMM operations. Furthermore, we introduce a dual-phase Curriculum Learning pipeline (Brain Stem Locking) utilizing Reinforcement Learning (RL) with Straight-Through Estimators (STE) and Elastic Weight Consolidation (EWC) to freeze foundational logical weights. The finalized architecture reduces memory constraints by approximately 80%, enabling 70B-parameter class models to run on single consumer-grade GPUs with near-zero latency mathematical precision.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.