LLM‐Powered Data Synthesis on Blockchain for Fusion Isomerism Learning in Heterogeneous Federated Systems

ABSTRACT Severe data starvation, architectural heterogeneity and Byzantine vulnerabilities fundamentally impede the deployment of robust multiclass classification models in decentralised edge environments. To address these intertwined challenges, we propose fusion isomerism learning (FusionIL) , a secure and domain‐agnostic framework that synergistically orchestrates isomerism large language models (IsoLLMs), consortium blockchain technology and heterogeneous federated learning. Unlike conventional approaches, FusionIL leverages a trusted committee of permissioned nodes to generate high‐fidelity synthetic data via feature‐driven textual prompts, forming IsoLLMs that capture structural heterogeneity across models. To prevent malicious data pollution, these generated samples undergo rigorous quantitative validation and are securely anchored via the practical Byzantine fault tolerance (PBFT) consensus mechanism. Furthermore, to mitigate the risk of generative model collapse, permissionless edge clients construct “synthetic interleaved datasets” by merging these verified synthetic samples with their private data under strict minority proportions. Subsequently, the framework employs the decentralised heterogeneous integration algorithm (DHIA) to mathematically reconcile and aggregate gradient updates from structurally diverse local models. This architectural innovation effectively eliminates the rigid homogeneity constraints of traditional federated networks. Extensive empirical evaluations across multiple benchmark datasets demonstrate that the synthetic interleaving strategy significantly enhances data diversity and global convergence. Consequently, FusionIL achieves competitive or leading performance against state‐of‐the‐art centralized and federated baselines in classification accuracy and system resilience, establishing a robust paradigm for secure, heterogeneous collaborative learning.

Authors

Institutions

Publication Details

Journal
CAAI Transactions on Intelligence Technology
Published
2026-10-04
DOI
https://doi.org/10.1049/cit2.70186
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

LLM‐Powered Data Synthesis on Blockchain for Fusion Isomerism Learning in Heterogeneous Federated Systems

Chuanbin Zhang, Xiaoqiang Teng, Zhihao Hao, Guancheng Wang et al.
CAAI Transactions on Intelligence Technology
Privacy-Preserving Technologies in Data
article

LLM‐Powered Data Synthesis on Blockchain for Fusion Isomerism Learning in Heterogeneous Federated Systems

Chuanbin Zhang, Xiaoqiang Teng, Zhihao Hao, Guancheng Wang, Longbing Cao, Han Yu
article en

Abstract

ABSTRACT Severe data starvation, architectural heterogeneity and Byzantine vulnerabilities fundamentally impede the deployment of robust multiclass classification models in decentralised edge environments. To address these intertwined challenges, we propose fusion isomerism learning (FusionIL) , a secure and domain‐agnostic framework that synergistically orchestrates isomerism large language models (IsoLLMs), consortium blockchain technology and heterogeneous federated learning. Unlike conventional approaches, FusionIL leverages a trusted committee of permissioned nodes to generate high‐fidelity synthetic data via feature‐driven textual prompts, forming IsoLLMs that capture structural heterogeneity across models. To prevent malicious data pollution, these generated samples undergo rigorous quantitative validation and are securely anchored via the practical Byzantine fault tolerance (PBFT) consensus mechanism. Furthermore, to mitigate the risk of generative model collapse, permissionless edge clients construct “synthetic interleaved datasets” by merging these verified synthetic samples with their private data under strict minority proportions. Subsequently, the framework employs the decentralised heterogeneous integration algorithm (DHIA) to mathematically reconcile and aggregate gradient updates from structurally diverse local models. This architectural innovation effectively eliminates the rigid homogeneity constraints of traditional federated networks. Extensive empirical evaluations across multiple benchmark datasets demonstrate that the synthetic interleaving strategy significantly enhances data diversity and global convergence. Consequently, FusionIL achieves competitive or leading performance against state‐of‐the‐art centralized and federated baselines in classification accuracy and system resilience, establishing a robust paradigm for secure, heterogeneous collaborative learning.

CAAI Transactions on Intelligence Technology
Nanyang Technological University (SG), Beijing Technology and Business University (CN), University of Macau (MO), Beijing Academy of Artificial Intelligence (CN), Macquarie University (AU)
National Natural Science Foundation of China, Beijing Technology and Business University
Openalex Percentile: Top 11%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.