Tensor-parallel Qwen3-30B inference on Raspberry Pi 5 devices with unequal memory

This report studies CPU inference of a quantized Qwen3-30B-A3B-Instruct-2507 model on two Raspberry Pi 5 devices with 8 GB and 16 GB of memory. A modified Distributed Llama runtime partitions attention and expert matrix computations within each decoder layer in a 1:3 ratio. In an initial Gigabit Ethernet experiment, two parallel runs achieved 6.027 and 6.385 generated tokens per second, with a median of 6.206. Two controlled sequential runs on the same devices achieved a median of 4.301 tokens per second. This corresponds to a measured throughput difference of 44.3%, with the sequential baseline also containing an ordering handshake. Separate monitored question runs averaged 5.480 tokens per second over direct Gigabit Ethernet and 4.196 through a 100 Mbps router. These observations demonstrate a practical deployment on devices with unequal memory and expose sensitivity to network conditions. A subsequent three-device extension stores and validates weight shards locally, reducing bulk startup traffic. Its inference results are qualified by power faults. True 2.5 GbE performance remains unmeasured; a faster switch alone cannot exceed the Pis' built-in Gigabit port limit. This record contains the technical report, a machine-readable benchmark evidence extract, and a README. It documents an experimental deployment; it has not been peer reviewed. The complete modified runtime and raw monitoring streams are not included. The manuscript discloses AI assistance used in its preparation.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22737194
Primary Topic
Network Packet Processing and Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Tensor-parallel Qwen3-30B inference on Raspberry Pi 5 devices with unequal memory

Lalit Belwal
Zenodo (CERN European Organization for Nuclear Research)
Network Packet Processing and Optimization
article

Tensor-parallel Qwen3-30B inference on Raspberry Pi 5 devices with unequal memory

Lalit Belwal
article en

Abstract

This report studies CPU inference of a quantized Qwen3-30B-A3B-Instruct-2507 model on two Raspberry Pi 5 devices with 8 GB and 16 GB of memory. A modified Distributed Llama runtime partitions attention and expert matrix computations within each decoder layer in a 1:3 ratio. In an initial Gigabit Ethernet experiment, two parallel runs achieved 6.027 and 6.385 generated tokens per second, with a median of 6.206. Two controlled sequential runs on the same devices achieved a median of 4.301 tokens per second. This corresponds to a measured throughput difference of 44.3%, with the sequential baseline also containing an ordering handshake. Separate monitored question runs averaged 5.480 tokens per second over direct Gigabit Ethernet and 4.196 through a 100 Mbps router. These observations demonstrate a practical deployment on devices with unequal memory and expose sensitivity to network conditions. A subsequent three-device extension stores and validates weight shards locally, reducing bulk startup traffic. Its inference results are qualified by power faults. True 2.5 GbE performance remains unmeasured; a faster switch alone cannot exceed the Pis' built-in Gigabit port limit. This record contains the technical report, a machine-readable benchmark evidence extract, and a README. It documents an experimental deployment; it has not been peer reviewed. The complete modified runtime and raw monitoring streams are not included. The manuscript discloses AI assistance used in its preparation.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 5%
Network Packet Processing and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Tensor-parallel Qwen3-30B inference on Raspberry Pi 5 devices with unequal memory — Lalit Belwal · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS