Tensor-parallel Qwen3-30B inference on Raspberry Pi 5 devices with unequal memory
This report studies CPU inference of a quantized Qwen3-30B-A3B-Instruct-2507 model on two Raspberry Pi 5 devices with 8 GB and 16 GB of memory. A modified Distributed Llama runtime partitions attention and expert matrix computations within each decoder layer in a 1:3 ratio. In an initial Gigabit Ethernet experiment, two parallel runs achieved 6.027 and 6.385 generated tokens per second, with a median of 6.206. Two controlled sequential runs on the same devices achieved a median of 4.301 tokens per second. This corresponds to a measured throughput difference of 44.3%, with the sequential baseline also containing an ordering handshake. Separate monitored question runs averaged 5.480 tokens per second over direct Gigabit Ethernet and 4.196 through a 100 Mbps router. These observations demonstrate a practical deployment on devices with unequal memory and expose sensitivity to network conditions. A subsequent three-device extension stores and validates weight shards locally, reducing bulk startup traffic. Its inference results are qualified by power faults. True 2.5 GbE performance remains unmeasured; a faster switch alone cannot exceed the Pis' built-in Gigabit port limit. This record contains the technical report, a machine-readable benchmark evidence extract, and a README. It documents an experimental deployment; it has not been peer reviewed. The complete modified runtime and raw monitoring streams are not included. The manuscript discloses AI assistance used in its preparation.
Authors
- Lalit Belwal
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22737194
- Primary Topic
- Network Packet Processing and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00