Where the Cheapest Token Wins, Pays, and Fails: A Predictive, Stress-Tested Map of a Frozen Single-Token Edge Encoder

Where the Cheapest Token Wins, Pays, and Fails: A Predictive, Stress-Tested Map of a Frozen Single-Token Edge Encoder Randolph James Ferlic, M.D. and Kimberly Kate Ferlic — Fieldstone Analytics, LLC, Austin, TX, USA Preprint · Zenodo DOI: 10.5281/zenodo.22945419 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract The 2026 wave of edge artificial intelligence optimizes the expensive tier: new mobile silicon runs generative, agentic models on-device at unprecedented performance-per-watt. Almost nothing has been done to make the always-on sensing-and-decision tier — the layer that decides when the expensive tier should even wake — comparably cheap. We study a frozen, deterministic, class-discriminant single-token encoder (features → a supervised linear-discriminant ⊕ principal-component subspace → a k-means codebook of at most 256 cells → an 8-bit token → a per-cell lookup-table decision) as a candidate Tier-0 layer that sits beneath the neural processing unit (NPU). One decision costs on the order of 20,000 operations — roughly four to five orders of magnitude fewer than a single generative token — and runs on a sensor hub, digital signal processor, or microcontroller with no NPU; priced in published per-operation energies, its arithmetic is tens of nanojoules or less per decision, roughly two orders of magnitude leaner than a deployed keyword-spotting wake-up. Rather than argue merely that such an encoder "works," we build and adversarially stress a predictive map of where it wins, pays a cost, and fails. Across ten physically distinct public datasets and a same-frozen pipeline, two axes emerge. On the accuracy axis, the token is competitive-to-winning against a strong gradient-boosted model where discriminative information lives in per-channel time-frequency statistics (bearing vibration: +0.16 to +0.21 AUROC; surgical kinematics: +0.067; electrocardiography: within +0.006), pays a small-to-moderate cost tax where the signal is per-channel but harder (surface electromyography ≈ 0.05–0.10; financial volatility-regime detection ≈ 0.12), and fails where the discriminative information is fundamentally spatial (electroencephalographic motor imagery: near chance, features discard the spatial code). On the personalization axis, a label-free per-entity "self-twin" recalibration recovers drift that is a local-distribution shift (sensor re-donning, cross-subject, cross-machine-load, cross-surgeon: full or 70–88% recovery) but fails where the non-stationarity is stationary-but-noisy (financial regime "drift," where more history beats recent data). We subject every load-bearing claim to four pre-registered adversarial rounds (twenty-plus attacks), fold each concession, and anchor the map with two genuine negatives (spatial-signal failure; self-twin failure). We further reconcile an apparent contradiction — that electrocardiography is simultaneously a "cheapest digital twin" showcase and a "cost-tax" domain — by showing the two claims measure different axes and are consistent. The token is, and remains, a cost and label-efficiency instrument, not an accuracy champion; its value on the agentic edge is a sub-milliwatt, auditable, private, self-healing Tier-0 layer whose behavior is now predictable from a domain's signal structure. This is a characterization of previously described, filed methods; it discloses no new algorithmic subject matter, and the per-deployment selection of configuration is retained as trade secret. Highlights · The reframe — a Tier-0 layer beneath the NPU: the 2026 silicon makes the expensive generative tier cheap; this positions the frozen token as the always-on tier below it — one decision ≈ 20,000 operations, roughly 10⁴–10⁵× fewer than a single generative token (tens of nanojoules or less per decision when priced in published per-operation energies, ~2 orders of magnitude leaner than a deployed keyword-spotting wake-up), on a sensor hub with no NPU, one byte on the wire. · The deliverable is a PREDICTIVE map, not a demo: a domain's signal structure predicts the token's behavior. It wins where information is per-channel time-frequency (bearing vibration +0.16/+0.21, surgical kinematics +0.067), is near-parity on electrocardiography (+0.006, the smallest tax), pays a bounded tax where per-channel but harder (sEMG ≈ 0.05–0.10, financial 0.117), and fails where the signal is fundamentally spatial (EEG motor imagery ≈ chance). · A label-free self-twin recovers LOCAL-distribution drift: cross-day surface electromyography (70–88%), cross-load bearing (full), cross-surgeon surgical (full), and per-patient electrocardiography (94% of a supervised ceiling without labels) — but fails on stationary-but-noisy financial regime non-stationarity, where more history beats recency. · Two genuine negatives anchor the map: a spatial-signal failure (EEG motor imagery, where a strong model on the same per-channel features fails too — the feature front-end, not the token) and a self-twin failure (financial). Honesty as evidence: a characterization that never reports a failure has not been stressed. · The target clinical population, tested: on eleven trans-radial amputees the encoder decodes intent (~4–6× chance), de-risking technology fit — while honestly bounding it (≈2.2× harder, a doubled tax, a useless cross-amputee population codebook, and label-hungry per-individual personalization). · Free-dial stack + multiplexing: an edge-default M0+D6+D1 stack (runtime-free/near-free, best label-efficiency, a built-in gate confidence), and one sub-milliwatt core driving up to K=40 heterogeneous detectors at ≈ 0 accuracy tax, with per-decision compute and emitted bits both falling as ≈1/K. · Four pre-registered adversarial rounds (20+ attacks): every load-bearing claim defended or conceded verbatim, including a methodology-leakage audit that exonerates the headline (rep/subject-disjoint) results; the ECG two-axis reconciliation in a single table. · All characterization of filed / published methods — no new algorithmic subject matter; the per-deployment configuration-selection procedure is a trade secret. What this record contains · `Manuscript_Paper49.pdf` — the manuscript, seven figures embedded (the two-tier architecture; the per-dial edge cost tiers; the K-detector multiplexing economics; the predictive two-axis map; the self-twin recovery across modalities; the token-versus-full-model bars; the electrocardiography two-axis), eight tables, and 73 references; and `Manuscript_Paper49.docx`, the editable source. · `PAPER_49_ZENODO_ARCHIVE.zip` — the reproducibility archive (md5 in ARCHIVE_MD5.txt): the frozen pre-registrations (the free-dial stack, the four data-in-hand adversarial rounds, the five new-dataset stresses, the ECG two-axis, and the same-day / multi-day self-twin pre-regs), the runners (the Tier-0 op-count and dial/multiplexing economics, the free-dial- stack validators, the four adversarial-round runners, the new-dataset stress runners, the ECG two-axis runner, the two self-twin runners, and the figure builders), the frozen token encoder module, the per-experiment result records (JSON) and per-simulation result notes (markdown), the seven figures, the manuscript source, and a README. All datasets are public and not redistributed (fetched from their public sources at run time); all paths and identifiers are scrubbed (absolute paths → PATH_TO_DATA/PATH_TO_SCRATCH, any cloud handle → MODAL_USER) and leak-scanned. Cite as R. J. Ferlic and K. K. Ferlic, "Where the cheapest token wins, pays, and fails: a predictive, stress-tested map of a frozen single-token edge encoder," Zenodo, 2026, doi: 10.5281/zenodo.22945419. License and patent notice Released under CC-BY 4.0. Consistent with that license, no patent or IP right of the authors is licensed, waived, or conveyed by this deposit. This work characterizes previously-described methods and discloses no new algorithmic subject matter; gradient boosting, k-means / vector quantization, product quantization, Fisher discriminant analysis, empirical-Bayes shrinkage, temperature scaling, and nearest-centroid novelty detection are established prior art, used only as tools. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, the multi-token / ladder and soft readouts, inference-time fusion, the foundation codebook, and the per-entity self-twin — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), the personalization / on-device-adaptation applications (priority U.S. Application No. 19/467,303 and its continuations), the multi-token / ladder application (No. 64/119,487), and the inference-time fusion application (No. 64/137,805). The per-deployment selection of configuration is retained as a trade secret and is not disclosed here. © 2026 Fieldstone Analytics, LLC and the authors. Inquiries: [email protected]. Companion deposits (spiral-domain-encoder-campaign) · The Configurable Bottleneck (the platform whose dials are edge-characterized here): doi:10.5281/zenodo.22884023 · The Cheapest Digital Twin (the self-twin whose scope this maps): doi:10.5281/zenodo.22922678 · Class-discriminant codebook construction (the base encoder): doi:10.5281/zenodo.20788187 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · Label-free inference-time channel fusion: doi:10.5281/zenodo.22046713 · The predictive reach of a decision token: doi:10.5281/zenodo.22736921 · Non-invertible but not anonymous (privacy): doi:10.5281/zenodo.22819210 · Unlinkable but not anonymous (privacy): doi:10.5281/zenodo.22838120 · The price of the bottleneck (deployment): doi:10.5281/zenodo.22838118 · Paying down the price of the bottleneck: doi:10.5281/zenodo.22866

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-24
DOI
https://doi.org/10.5281/zenodo.22945419
Primary Topic
Advancements in Semiconductor Devices and Circuit Design
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Where the Cheapest Token Wins, Pays, and Fails: A Predictive, Stress-Tested Map of a Frozen Single-Token Edge Encoder

Randolph James Ferlic, Kimberly Kate Ferlic
Zenodo (CERN European Organization for Nuclear Research)
Advancements in Semiconductor Devices and Circuit Design
preprint

Where the Cheapest Token Wins, Pays, and Fails: A Predictive, Stress-Tested Map of a Frozen Single-Token Edge Encoder

Randolph James Ferlic, Kimberly Kate Ferlic
preprint en

Abstract

Where the Cheapest Token Wins, Pays, and Fails: A Predictive, Stress-Tested Map of a Frozen Single-Token Edge Encoder Randolph James Ferlic, M.D. and Kimberly Kate Ferlic — Fieldstone Analytics, LLC, Austin, TX, USA Preprint · Zenodo DOI: 10.5281/zenodo.22945419 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract The 2026 wave of edge artificial intelligence optimizes the expensive tier: new mobile silicon runs generative, agentic models on-device at unprecedented performance-per-watt. Almost nothing has been done to make the always-on sensing-and-decision tier — the layer that decides when the expensive tier should even wake — comparably cheap. We study a frozen, deterministic, class-discriminant single-token encoder (features → a supervised linear-discriminant ⊕ principal-component subspace → a k-means codebook of at most 256 cells → an 8-bit token → a per-cell lookup-table decision) as a candidate Tier-0 layer that sits beneath the neural processing unit (NPU). One decision costs on the order of 20,000 operations — roughly four to five orders of magnitude fewer than a single generative token — and runs on a sensor hub, digital signal processor, or microcontroller with no NPU; priced in published per-operation energies, its arithmetic is tens of nanojoules or less per decision, roughly two orders of magnitude leaner than a deployed keyword-spotting wake-up. Rather than argue merely that such an encoder "works," we build and adversarially stress a predictive map of where it wins, pays a cost, and fails. Across ten physically distinct public datasets and a same-frozen pipeline, two axes emerge. On the accuracy axis, the token is competitive-to-winning against a strong gradient-boosted model where discriminative information lives in per-channel time-frequency statistics (bearing vibration: +0.16 to +0.21 AUROC; surgical kinematics: +0.067; electrocardiography: within +0.006), pays a small-to-moderate cost tax where the signal is per-channel but harder (surface electromyography ≈ 0.05–0.10; financial volatility-regime detection ≈ 0.12), and fails where the discriminative information is fundamentally spatial (electroencephalographic motor imagery: near chance, features discard the spatial code). On the personalization axis, a label-free per-entity "self-twin" recalibration recovers drift that is a local-distribution shift (sensor re-donning, cross-subject, cross-machine-load, cross-surgeon: full or 70–88% recovery) but fails where the non-stationarity is stationary-but-noisy (financial regime "drift," where more history beats recent data). We subject every load-bearing claim to four pre-registered adversarial rounds (twenty-plus attacks), fold each concession, and anchor the map with two genuine negatives (spatial-signal failure; self-twin failure). We further reconcile an apparent contradiction — that electrocardiography is simultaneously a "cheapest digital twin" showcase and a "cost-tax" domain — by showing the two claims measure different axes and are consistent. The token is, and remains, a cost and label-efficiency instrument, not an accuracy champion; its value on the agentic edge is a sub-milliwatt, auditable, private, self-healing Tier-0 layer whose behavior is now predictable from a domain's signal structure. This is a characterization of previously described, filed methods; it discloses no new algorithmic subject matter, and the per-deployment selection of configuration is retained as trade secret. Highlights · The reframe — a Tier-0 layer beneath the NPU: the 2026 silicon makes the expensive generative tier cheap; this positions the frozen token as the always-on tier below it — one decision ≈ 20,000 operations, roughly 10⁴–10⁵× fewer than a single generative token (tens of nanojoules or less per decision when priced in published per-operation energies, ~2 orders of magnitude leaner than a deployed keyword-spotting wake-up), on a sensor hub with no NPU, one byte on the wire. · The deliverable is a PREDICTIVE map, not a demo: a domain's signal structure predicts the token's behavior. It wins where information is per-channel time-frequency (bearing vibration +0.16/+0.21, surgical kinematics +0.067), is near-parity on electrocardiography (+0.006, the smallest tax), pays a bounded tax where per-channel but harder (sEMG ≈ 0.05–0.10, financial 0.117), and fails where the signal is fundamentally spatial (EEG motor imagery ≈ chance). · A label-free self-twin recovers LOCAL-distribution drift: cross-day surface electromyography (70–88%), cross-load bearing (full), cross-surgeon surgical (full), and per-patient electrocardiography (94% of a supervised ceiling without labels) — but fails on stationary-but-noisy financial regime non-stationarity, where more history beats recency. · Two genuine negatives anchor the map: a spatial-signal failure (EEG motor imagery, where a strong model on the same per-channel features fails too — the feature front-end, not the token) and a self-twin failure (financial). Honesty as evidence: a characterization that never reports a failure has not been stressed. · The target clinical population, tested: on eleven trans-radial amputees the encoder decodes intent (~4–6× chance), de-risking technology fit — while honestly bounding it (≈2.2× harder, a doubled tax, a useless cross-amputee population codebook, and label-hungry per-individual personalization). · Free-dial stack + multiplexing: an edge-default M0+D6+D1 stack (runtime-free/near-free, best label-efficiency, a built-in gate confidence), and one sub-milliwatt core driving up to K=40 heterogeneous detectors at ≈ 0 accuracy tax, with per-decision compute and emitted bits both falling as ≈1/K. · Four pre-registered adversarial rounds (20+ attacks): every load-bearing claim defended or conceded verbatim, including a methodology-leakage audit that exonerates the headline (rep/subject-disjoint) results; the ECG two-axis reconciliation in a single table. · All characterization of filed / published methods — no new algorithmic subject matter; the per-deployment configuration-selection procedure is a trade secret. What this record contains · `Manuscript_Paper49.pdf` — the manuscript, seven figures embedded (the two-tier architecture; the per-dial edge cost tiers; the K-detector multiplexing economics; the predictive two-axis map; the self-twin recovery across modalities; the token-versus-full-model bars; the electrocardiography two-axis), eight tables, and 73 references; and `Manuscript_Paper49.docx`, the editable source. · `PAPER_49_ZENODO_ARCHIVE.zip` — the reproducibility archive (md5 in ARCHIVE_MD5.txt): the frozen pre-registrations (the free-dial stack, the four data-in-hand adversarial rounds, the five new-dataset stresses, the ECG two-axis, and the same-day / multi-day self-twin pre-regs), the runners (the Tier-0 op-count and dial/multiplexing economics, the free-dial- stack validators, the four adversarial-round runners, the new-dataset stress runners, the ECG two-axis runner, the two self-twin runners, and the figure builders), the frozen token encoder module, the per-experiment result records (JSON) and per-simulation result notes (markdown), the seven figures, the manuscript source, and a README. All datasets are public and not redistributed (fetched from their public sources at run time); all paths and identifiers are scrubbed (absolute paths → PATH_TO_DATA/PATH_TO_SCRATCH, any cloud handle → MODAL_USER) and leak-scanned. Cite as R. J. Ferlic and K. K. Ferlic, "Where the cheapest token wins, pays, and fails: a predictive, stress-tested map of a frozen single-token edge encoder," Zenodo, 2026, doi: 10.5281/zenodo.22945419. License and patent notice Released under CC-BY 4.0. Consistent with that license, no patent or IP right of the authors is licensed, waived, or conveyed by this deposit. This work characterizes previously-described methods and discloses no new algorithmic subject matter; gradient boosting, k-means / vector quantization, product quantization, Fisher discriminant analysis, empirical-Bayes shrinkage, temperature scaling, and nearest-centroid novelty detection are established prior art, used only as tools. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, the multi-token / ladder and soft readouts, inference-time fusion, the foundation codebook, and the per-entity self-twin — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder), the personalization / on-device-adaptation applications (priority U.S. Application No. 19/467,303 and its continuations), the multi-token / ladder application (No. 64/119,487), and the inference-time fusion application (No. 64/137,805). The per-deployment selection of configuration is retained as a trade secret and is not disclosed here. © 2026 Fieldstone Analytics, LLC and the authors. Inquiries: [email protected]. Companion deposits (spiral-domain-encoder-campaign) · The Configurable Bottleneck (the platform whose dials are edge-characterized here): doi:10.5281/zenodo.22884023 · The Cheapest Digital Twin (the self-twin whose scope this maps): doi:10.5281/zenodo.22922678 · Class-discriminant codebook construction (the base encoder): doi:10.5281/zenodo.20788187 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · Label-free inference-time channel fusion: doi:10.5281/zenodo.22046713 · The predictive reach of a decision token: doi:10.5281/zenodo.22736921 · Non-invertible but not anonymous (privacy): doi:10.5281/zenodo.22819210 · Unlinkable but not anonymous (privacy): doi:10.5281/zenodo.22838120 · The price of the bottleneck (deployment): doi:10.5281/zenodo.22838118 · Paying down the price of the bottleneck: doi:10.5281/zenodo.22866

Zenodo (CERN European Organization for Nuclear Research)
EP Analytics (United States) (US)
Reduced inequalities
Advancements in Semiconductor Devices and Circuit Design
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.