Interoperable Modular Language Models I: From Decomposition to Interchangeability
> Large language models are usually trained, distributed, and consumed as monolithic systems even though their internal computations are structurally modular. This paper asks a stronger question than model splitting, routing, merging, or adapter composition: can internal language-model subsystems become independently packaged, qualified, replaced, recombined, and transported across hosts in a way that resembles a technical component ecosystem? We study this question through a prospectively governed program spanning native modular training, retrofit decomposition, component-resolution selection, supplier-like replacement, cross-supplier recomposition, pre-assembly risk prediction, host portability, cross-architecture bridging, and a conditional industrial-organization model. The central result is a separation between decomposition and interchangeability. Executed package chains can preserve the behavior of two pinned pretrained models at tested block granularities, but independent replacement is substantially harder. Across three structural resolutions, no resolution dominates on quality, interaction, latency, and certification burden. Four frozen supplier-like training arms fail zero-pair-calibration transfer to held-out recipient hosts. A pre-assembly handshake predictor improves finite-library risk prediction over additive single-component effects, yet a complete 25-assembly recipient oracle contains no safe assembly, separating selector error from library insufficiency. Reusable per-entity representation registration reduces direct-swap loss by about 1.9-2.0 nats, but remains far outside frozen quality margins and is inferior to pair-specific calibration; pair-specific calibration also fails the absolute portability gate. Pair-specific linear Pythia/GPT-2 bridges execute but incur perplexity ratios of 2.04x and 6.43x. We therefore propose an evidence ladder that separates mechanical execution, representation compatibility, behavioral qualification, composition qualification, and portability. The paper does not claim that an LLM component market is currently demonstrated. Instead, it maps the technical gates between exact decomposition and genuine interchangeability and defines prospective research branches for learning interoperable interfaces, certifying component libraries, mediating architecture boundaries, and analyzing qualified component ecosystems.
Authors
- Minseong Sim (ORCID: https://orcid.org/0009-0002-3105-2877)
Institutions
- Hankuk University of Foreign Studies (KR)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23133107
- Primary Topic
- Topic Modeling
- Type
- preprint