Tensor-parallel Qwen3-30B inference on three Raspberry Pi devices: Gigabit scaling and synchronization evidence
This technical note extends an unequal-memory, two-device Qwen3 inference study to three Raspberry Pi devices connected by Gigabit Ethernet. Two Raspberry Pi 5 boards with 8 GB and 16 GB RAM and a 16 GB Compute Module 5 execute tensor partitions within every decoder layer of a requantized Qwen3-30B-A3B-Instruct-2507 model. Equal expert-channel partitions and persistent local weight shards replace the previous 25%/75% work split. Three complete three-device batches averaged 8.619, 8.267 and 8.377 generated tokens/s, compared with a fresh two-device parallel baseline of 5.737 tokens/s. The latest batch therefore shows 46.00% higher decode throughput for the same request and output count. A separate instrumented diagnostic attributed 28.15% of root timed steps to synchronization, including transfer and peer readiness. The latest API run reached 85.1°C on the Compute Module and observed its loaded CPU clock fall to 1.5 GHz; no model swap or undervoltage was sampled in that batch. TCP retransmissions persisted despite Gigabit links. These are deployment-level observations with thermal, memory, output-trace and scheduling confounders, rather than an isolated hardware scaling or three-device sequential comparison. Curated raw measurements and analysis scripts accompany the report. Version 2.0 extends the earlier two-Pi technical note with the three-Pi Gigabit deployment, all complete repeat batches, local-shard startup measurements, thermal/CPU/TCP monitoring and a separate compute/Sync diagnostic. The 46.00% result compares three parallel Pis against two parallel Pis; no three-Pi sequential speedup is claimed. The release contains a 7-page manuscript, benchmark-evidence.json, README.txt and reproducibility.zip with 63 curated source files, raw numeric monitoring streams, source hashes and scripts that regenerate the analysis. Operational identifiers are replaced with placeholders. Full model weights and a complete buildable modified inference runtime are not bundled. The manuscript discloses AI assistance and experimental limitations. This technical note has not been peer reviewed. Report, evidence and original analysis scripts: CC BY 4.0. The three Distributed Llama timing-reference source files retain their included upstream MIT license.
Authors
- Lalit Belwal
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22773703
- Primary Topic
- Network Packet Processing and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00