Multi-modal parameter-efficient fine-tuning via hypergraph adapters

Parameter-efficient fine-tuning (PEFT) is an attractive strategy for adapting vision–language models when compute, memory, or labeled data are limited. However, most multimodal PEFT methods still operate on individual samples or pairwise relations, which leaves higher-order structure shared across semantically related examples largely unexploited. We propose HGA-Net, a hypergraph adapter that constructs a mini-batch hypergraph in a joint image–text embedding space and performs lightweight hypergraph message passing inside an adapter bottleneck. To improve stability under noisy captions and ambiguous neighborhoods, we further introduce a soft-incidence formulation that replaces hard hyperedge membership with similarity-weighted participation. We evaluate the method in a scope-matched set-conditioned regime, where inference aggregates only local batch context, and we additionally study a fixed-reference variant for query-independent deployment. Under a unified frozen-backbone interface built on CLIP ViT-B/16 and shared text inputs across all compared modules, HGA-Net achieves top-1 accuracies of 99.71% on Flowers102 and 93.20% on Oxford-IIIT Pets while using 1.573M adapter parameters. The expanded experiments further extend the evaluation to eleven recognition benchmarks, multiple shot budgets (1, 2, 4, 8, and 16 shots), batch-size and batch-composition sensitivity, fixed-reference inference, latency, and a larger CLIP ViT-L/14 backbone. These results indicate that explicitly modeling higher-order cross-sample structure can be an effective inductive bias for set-conditioned multimodal adaptation.

Authors

Institutions

Publication Details

Journal
Complex & Intelligent Systems
Published
2026-10-05
DOI
https://doi.org/10.1007/s40747-026-02534-7
Primary Topic
Domain Adaptation and Few-Shot Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-modal parameter-efficient fine-tuning via hypergraph adapters

Jiaxuan Lu, Wenjie Pan, Dongmei Wang, Junyan Lv
Complex & Intelligent Systems
Domain Adaptation and Few-Shot Learning
article

Multi-modal parameter-efficient fine-tuning via hypergraph adapters

Jiaxuan Lu, Wenjie Pan, Dongmei Wang, Junyan Lv
article en

Abstract

Parameter-efficient fine-tuning (PEFT) is an attractive strategy for adapting vision–language models when compute, memory, or labeled data are limited. However, most multimodal PEFT methods still operate on individual samples or pairwise relations, which leaves higher-order structure shared across semantically related examples largely unexploited. We propose HGA-Net, a hypergraph adapter that constructs a mini-batch hypergraph in a joint image–text embedding space and performs lightweight hypergraph message passing inside an adapter bottleneck. To improve stability under noisy captions and ambiguous neighborhoods, we further introduce a soft-incidence formulation that replaces hard hyperedge membership with similarity-weighted participation. We evaluate the method in a scope-matched set-conditioned regime, where inference aggregates only local batch context, and we additionally study a fixed-reference variant for query-independent deployment. Under a unified frozen-backbone interface built on CLIP ViT-B/16 and shared text inputs across all compared modules, HGA-Net achieves top-1 accuracies of 99.71% on Flowers102 and 93.20% on Oxford-IIIT Pets while using 1.573M adapter parameters. The expanded experiments further extend the evaluation to eleven recognition benchmarks, multiple shot budgets (1, 2, 4, 8, and 16 shots), batch-size and batch-composition sensitivity, fixed-reference inference, latency, and a larger CLIP ViT-L/14 backbone. These results indicate that explicitly modeling higher-order cross-sample structure can be an effective inductive bias for set-conditioned multimodal adaptation.

Complex & Intelligent Systems
Shanghai Artificial Intelligence Laboratory (CN), Nanjing City Vocational College (CN)
Openalex Percentile: Top 10%
Domain Adaptation and Few-Shot Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multi-modal parameter-efficient fine-tuning via hypergraph adapters — Jiaxuan Lu, Wenjie Pan, et al. · Complex & Intelligent Systems (2026) | TGRS Research Map | TGRS