Semantic Fragmentation and Stochastic Assembly: A Protocol for Decentralized Language-Model Inference over Untrusted Volunteer Nodes

Peer-to-peer language-model inference has so far been pursued by splitting the model: transformer layers or tensors are distributed across machines, and intermediate activations traverse the public internet on every generated token. This places the design squarely against a bandwidth gap of roughly five orders of magnitude and a latency gap of four to five, between datacenter interconnect and consumer last-mile links. This paper specifies Swarmbly, a protocol that distributes the problem instead. A client-side orchestrator, itself a small language model, decomposes a request into a directed acyclic graph of semantic micro-tasks; each micro-task is dispatched once, asynchronously, to a volunteer node running a complete small model (1–8B parameters); returned fragments — contigs, in the genome-assembly vocabulary the design borrows — are verified, selected and spliced locally. Network traversal occurs once per fragment per session rather than once per layer per token. Swarmbly AI A decentralized inference protocol that fragments the problem, not the model. Swarmbly dispatches semantic micro-tasks to volunteer nodes running complete small language models (SLMs, 1–8B), and reassembles the answers on the client with an orchestrator SLM — using genome shotgun assembly (reads, contigs, overlap, scaffolding, consensus) as its design vocabulary. Existing peer-to-peer inference systems split the model: layers or tensors live on different machines, and activations cross the public internet on every token. Swarmbly splits the problem: each fragment crosses the network once, and every worker runs a whole, small, independent model.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23031305
Primary Topic
Advanced Graph Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Semantic Fragmentation and Stochastic Assembly: A Protocol for Decentralized Language-Model Inference over Untrusted Volunteer Nodes

Sebastian A. Espinoza‐Ulloa
Zenodo (CERN European Organization for Nuclear Research)
Advanced Graph Neural Networks
article

Semantic Fragmentation and Stochastic Assembly: A Protocol for Decentralized Language-Model Inference over Untrusted Volunteer Nodes

Sebastian A. Espinoza‐Ulloa
article en

Abstract

Peer-to-peer language-model inference has so far been pursued by splitting the model: transformer layers or tensors are distributed across machines, and intermediate activations traverse the public internet on every generated token. This places the design squarely against a bandwidth gap of roughly five orders of magnitude and a latency gap of four to five, between datacenter interconnect and consumer last-mile links. This paper specifies Swarmbly, a protocol that distributes the problem instead. A client-side orchestrator, itself a small language model, decomposes a request into a directed acyclic graph of semantic micro-tasks; each micro-task is dispatched once, asynchronously, to a volunteer node running a complete small model (1–8B parameters); returned fragments — contigs, in the genome-assembly vocabulary the design borrows — are verified, selected and spliced locally. Network traversal occurs once per fragment per session rather than once per layer per token. Swarmbly AI A decentralized inference protocol that fragments the problem, not the model. Swarmbly dispatches semantic micro-tasks to volunteer nodes running complete small language models (SLMs, 1–8B), and reassembles the answers on the client with an orchestrator SLM — using genome shotgun assembly (reads, contigs, overlap, scaffolding, consensus) as its design vocabulary. Existing peer-to-peer inference systems split the model: layers or tensors live on different machines, and activations cross the public internet on every token. Swarmbly splits the problem: each fragment crosses the network once, and every worker runs a whole, small, independent model.

Zenodo (CERN European Organization for Nuclear Research)
Pontificia Universidad Católica del Ecuador (EC)
Openalex Percentile: Top 9%
Advanced Graph Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Semantic Fragmentation and Stochastic Assembly: A Protocol for Decentralized Language-Model Inference over Untrusted Volunteer Nodes — Sebastian A. Espinoza‐Ulloa · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS