Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform

Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified GDDR6 memory and a Vulkan-only GPU compute path), as a single-board testbed. The primary contribution is a reproducible characterisation of that platform, released with the artefacts needed to repeat it. Three observations are not specific to the board. Two are benchmark-validity cautions: (i) a silent prompt-truncation failure mode in the Ollama runtime, which affects filled-context benchmarking unless the API's prompt_eval_count field is verified per request, and (ii) the distinction between context allocation and context utilisation, which can produce headline "128K context" numbers backed by a much shorter prompt. The third (iii) is confirmatory: sparse mixture-of-experts (MoE) generation throughput tracks the active-parameter count when matrix accelerators are absent, observed here on hardware where, to the author's knowledge, it had not been measured before. A per-token byte accounting from the model tensor tables predicts the observed separation from a same-size dense comparator at matched quantisation to within 4 %, and the deployable edge over 14 B dense models is a modest ≈14–25 %. The evidence rests on a 31-model cohort (3–35 B parameters) and covers generation speed (run-to-run within-cell coefficient of variation ≤0.5 % across the canonical n=3 cells under a pinned governor), filled-context scaling with real-token payloads, a five-task competence-preservation probe (a check that a configuration is not broken, not a capability score), cold-start latency, and clock and thermal telemetry. Three runtime settings the results depend on were measured directly. The 100-token decode window and flash attention hold. Quantising the KV cache to 4 bits costs up to 1.6 % perplexity on most models, more on several others, and collapses the small Qwen2-architecture builds through their key cache, whereas an 8-bit cache stays within ±0.5 % of FP16. The deployment recipe, a kernel TTM (Translation Table Manager) pages_limit adjustment and an FP-validated reuse of a community 40-CU core unlock, is documented as the enabling infrastructure. All measurements come from a single board, driver, and firmware; the scope limits this implies are stated in the threats analysis. Version 3 is the accepted manuscript of Journal of Universal Computer Science submission #204239 (accepted 3 October 2026, scheduled for J.UCS 33(6), June 2027), in the form submitted for copy editing. Version 2 was the manuscript as submitted for review on 16 June 2026. The measurement artefacts (harnesses, raw results, supplementary material) are archived separately at 10.5281/zenodo.22668983.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23157622
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform

Artur Andrzejczak
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

Deploying LLM Inference on a Repurposed UMA APU: Transferable Lessons from a Vulkan-Only, 16 GB Edge Platform

Artur Andrzejczak
preprint en

Abstract

Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified GDDR6 memory and a Vulkan-only GPU compute path), as a single-board testbed. The primary contribution is a reproducible characterisation of that platform, released with the artefacts needed to repeat it. Three observations are not specific to the board. Two are benchmark-validity cautions: (i) a silent prompt-truncation failure mode in the Ollama runtime, which affects filled-context benchmarking unless the API's prompt_eval_count field is verified per request, and (ii) the distinction between context allocation and context utilisation, which can produce headline "128K context" numbers backed by a much shorter prompt. The third (iii) is confirmatory: sparse mixture-of-experts (MoE) generation throughput tracks the active-parameter count when matrix accelerators are absent, observed here on hardware where, to the author's knowledge, it had not been measured before. A per-token byte accounting from the model tensor tables predicts the observed separation from a same-size dense comparator at matched quantisation to within 4 %, and the deployable edge over 14 B dense models is a modest ≈14–25 %. The evidence rests on a 31-model cohort (3–35 B parameters) and covers generation speed (run-to-run within-cell coefficient of variation ≤0.5 % across the canonical n=3 cells under a pinned governor), filled-context scaling with real-token payloads, a five-task competence-preservation probe (a check that a configuration is not broken, not a capability score), cold-start latency, and clock and thermal telemetry. Three runtime settings the results depend on were measured directly. The 100-token decode window and flash attention hold. Quantising the KV cache to 4 bits costs up to 1.6 % perplexity on most models, more on several others, and collapses the small Qwen2-architecture builds through their key cache, whereas an 8-bit cache stays within ±0.5 % of FP16. The deployment recipe, a kernel TTM (Translation Table Manager) pages_limit adjustment and an FP-validated reuse of a community 40-CU core unlock, is documented as the enabling infrastructure. All measurements come from a single board, driver, and firmware; the scope limits this implies are stated in the threats analysis. Version 3 is the accepted manuscript of Journal of Universal Computer Science submission #204239 (accepted 3 October 2026, scheduled for J.UCS 33(6), June 2027), in the form submitted for copy editing. Version 2 was the manuscript as submitted for review on 16 June 2026. The measurement artefacts (harnesses, raw results, supplementary material) are archived separately at 10.5281/zenodo.22668983.

Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.