Cartridge-Oriented Architecture for Bare-Metal AI Inference on Resource-Constrained Edge Devices
Resource-constrained edge devices increasingly require local AI inference while providing limited memory, storage, compute capability, and operating-system support. This paper presents a cartridge-oriented architecture in which model-specific inference information is packaged as a compact, validated artifact executed by a reusable native inference engine. Unlike conventional deployments that embed models within a general-purpose inference runtime, the proposed approach establishes an explicit deployment boundary between the reusable execution engine and replaceable AI cartridges. A prototype was implemented as a bare-metal system and evaluated using QEMU RISC-V and ARM Cortex-M3 targets. The experiments demonstrate independent artifact replacement, loader-level validation, execution of multiple operator structures, and byte-identical cartridge portability across the two architectures. The results provide evidence that a model-specific deployment artifact can be separated from a reusable native execution engine without requiring a general-purpose operating system or interpreter. The idea is conceptually inspired by the cartridge-based architecture of classic Mario game consoles: a reusable console provides the execution platform, while a replaceable cartridge contains game-specific content and logic. Note: This is an unpublished preprint version of a manuscript that was previously considered at an IEEE venue. The authors retain full copyright of this early draft version. The work is being self-archived here prior to extension and future journal/conference submission.
Authors
- Abhinandan Bhadauria
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23125459
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint