A Unified Intermediate Representation and Execution Architecture for Heterogeneous and Distributed Neural Simulation
Equation-based neural modeling offers flexibility, but executing models across processors and nodes requires explicit control of order, ownership, and numerical behavior. We present brian2-atlas, a rebuilt execution architecture retaining Brian2's modeling interface and bringing heterogeneous and distributed targets under a common model-to-execution contract. Its unified intermediate representation, AtlasIR (Atlas Intermediate Representation), describes typed state, clocks, schedules, effects, events, and random-stream identities, and is checked by an independent Rust validator. Execution-plan analysis constrains physical transformations; target-specific implementations support Rust CPU execution, CUDA and Apple Metal GPUs, browser WebAssembly, and MPI within declared contracts. Ownership-based event processing, compact topology construction, and partitioned storage connect model preparation to distributed execution. Supported MPI pathways additionally preserve mutable synaptic state, delayed events, and boundary checkpoints across simulation segments. In a repeated full-connectome CPU cohort, the best measured Rust configuration reduced simulation-and-recording time by 5.43-fold relative to the best measured Brian2 C++ configuration on the same Linux host. A 298.9-million-synapse model completed on a 16 GB laptop while the compared C++ preparation path exceeded its controlled memory budget. Distributed execution completed 100.5 seconds of model time for 4,129,924 neurons and 24,126,516,728 recurrent synapses on four nodes and 32 ranks. A five-point synthetic weak-scaling study reached 860 million neurons and 860 billion recurrent connections for 100 ms on 30 hosts and 240 physical worker cores, increasing from 86 million neurons on three hosts. These establish tested capacity configurations rather than statistical equivalence or a matched speed advantage over NEST. Published-model workflows extend validation to explicit NMDA dynamics and contextual dendritic plasticity. A matched eight-core NMDA cohort achieves up to 2.10-fold acceleration, while a common-source dendritic-network retest achieves 1.28-fold acceleration after effect-safe endpoint-expression reuse; historical dendritic cohorts retain their separate source identities. Browser studies demonstrate portable reference execution and a separate full-connectome application. A bounded native-training extension supplies explicit forward and reverse execution plans with surrogate spike gradients. The results characterize semantic, numerical, and resource boundaries rather than universal compatibility or performance superiority. English preprint, version v1. This manuscript has not been peer reviewed. Related software archive: https://doi.org/10.5281/zenodo.23269098Related evidence archive: https://doi.org/10.5281/zenodo.23269145The software and evidence archive files are currently restricted; their metadata remain public.
Authors
- Xinjun Li
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-09
- DOI
- https://doi.org/10.5281/zenodo.23269879
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint