CARE-MoE: Conflict-Aware Routing-Guided Expert Mixture-of-Experts for Continual Learning

Continual learning (CL) aims to learn new tasks sequentially while retaining prior knowledge. Although recent mixture-of-experts (MoE) architectures mitigate catastrophic forgetting by reducing global parameter sharing, they introduce a new bottleneck: expert-level interference. Because multiple tasks inevitably share a subset of experts, these shared parameters are repeatedly exposed to heterogeneous task updates, which may contribute to the degradation of previously encoded knowledge. To tackle this, we propose CARE-MoE (Conflict-Aware Routing-Guided Expert Mixture-of-Experts), a unified framework designed to limit repeated expert sharing, characterize expert utilization, and regulate expert-level updates. First, CARE-MoE aims to limit excessive expert sharing via dual-granularity routing, which leverages both task-level structural and input-level semantic information to distribute representations effectively. Second, rather than using routing merely for expert selection, it diagnoses highly exposed experts via multi-dimensional routing signal aggregation, quantifying expert importance through selection frequency, cross-task overlap, and routing consistency. Finally, conflict-aware expert stabilization translates this diagnosis into targeted optimization, selectively directing preservation updates to the routing parameters associated with heavily shared experts and modulating their strength according to the directional agreement between adaptation and preservation, while allowing task-specific experts to flexibly adapt. Extensive experiments on standard and long-sequence CL benchmarks show that CARE-MoE achieves higher overall performance than the results reported for existing methods. In-depth analyses further indicate that our framework bridges routing behavior with parameter optimization, achieving a favorable balance between stability and plasticity in long-term continual learning.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-30
DOI
https://doi.org/10.3390/math14193557
Primary Topic
Domain Adaptation and Few-Shot Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CARE-MoE: Conflict-Aware Routing-Guided Expert Mixture-of-Experts for Continual Learning

Daeho Kim, SoYeop Yoo, Ok‐Ran Jeong, SeonGyoung Lee
Mathematics
Domain Adaptation and Few-Shot Learning
article

CARE-MoE: Conflict-Aware Routing-Guided Expert Mixture-of-Experts for Continual Learning

Daeho Kim, SoYeop Yoo, Ok‐Ran Jeong, SeonGyoung Lee
article en

Abstract

Continual learning (CL) aims to learn new tasks sequentially while retaining prior knowledge. Although recent mixture-of-experts (MoE) architectures mitigate catastrophic forgetting by reducing global parameter sharing, they introduce a new bottleneck: expert-level interference. Because multiple tasks inevitably share a subset of experts, these shared parameters are repeatedly exposed to heterogeneous task updates, which may contribute to the degradation of previously encoded knowledge. To tackle this, we propose CARE-MoE (Conflict-Aware Routing-Guided Expert Mixture-of-Experts), a unified framework designed to limit repeated expert sharing, characterize expert utilization, and regulate expert-level updates. First, CARE-MoE aims to limit excessive expert sharing via dual-granularity routing, which leverages both task-level structural and input-level semantic information to distribute representations effectively. Second, rather than using routing merely for expert selection, it diagnoses highly exposed experts via multi-dimensional routing signal aggregation, quantifying expert importance through selection frequency, cross-task overlap, and routing consistency. Finally, conflict-aware expert stabilization translates this diagnosis into targeted optimization, selectively directing preservation updates to the routing parameters associated with heavily shared experts and modulating their strength according to the directional agreement between adaptation and preservation, while allowing task-specific experts to flexibly adapt. Extensive experiments on standard and long-sequence CL benchmarks show that CARE-MoE achieves higher overall performance than the results reported for existing methods. In-depth analyses further indicate that our framework bridges routing behavior with parameter optimization, achieving a favorable balance between stability and plasticity in long-term continual learning.

MathematicsVol. 14(19)
Gachon University (KR), National Institute of Advanced Industrial Science and Technology (JP)
Openalex Percentile: Top 9%
Domain Adaptation and Few-Shot Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.