Efficient Multimodal Instruction Tuning Through Iterative Data Selection and Augmentation Under Fixed Data Budgets

Efficient multimodal instruction tuning requires reducing both training data volume and computational cost without substantially sacrificing downstream performance. Existing data selection methods can remove redundancy but cannot repair imperfect supervision or supplement under-represented capabilities. We propose SANet, a fixed-budget data construction framework that alternates Multi-Criteria Selection and Bi-Level Augmentation. SANet refines imperfect instruction–response pairs, generates samples for under-represented capabilities, and re-evaluates augmented candidates through subsequent selection before constructing the final subset. Using only 6% of the original training data, SANet achieves an average score of 61.88 over the three predefined primary benchmarks (ScienceQA, MMBench, and TextVQA), outperforming 6% Random Selection by 1.83 points on this primary metric. On the additional diagnostic GQA benchmark, however, SANet performs worse than Random Selection, revealing a limitation in fine-grained compositional reasoning. The complete SANet pipeline requires approximately 35 h under the same hardware configuration, including a non-negligible one-time subset-construction cost of approximately 29 h and 6 h for downstream fine-tuning, compared with 67 h for full-data fine-tuning. Because the constructed subset can be reused across different backbones, the one-time construction overhead can be amortized when multiple downstream models are fine-tuned. The constructed subset can also be directly reused without backbone-specific reconstruction, achieving average scores of 69.03 on Qwen2-VL-7B and 87.56 on InternVL2.5-8B. These results indicate that SANet provides a data-efficient multimodal instruction-tuning pipeline with practical cross-model reusability. The reported multi-seed variation is based on three downstream fine-tuning runs with fixed constructed subsets and therefore reflects training-seed sensitivity rather than full-pipeline statistical uncertainty.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-28
DOI
https://doi.org/10.3390/electronics15194461
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Efficient Multimodal Instruction Tuning Through Iterative Data Selection and Augmentation Under Fixed Data Budgets

Yi Wang, H Y Lin, Yali Wang, Yinan He
Electronics
Natural Language Processing Techniques
article

Efficient Multimodal Instruction Tuning Through Iterative Data Selection and Augmentation Under Fixed Data Budgets

Yi Wang, H Y Lin, Yali Wang, Yinan He
article en

Abstract

Efficient multimodal instruction tuning requires reducing both training data volume and computational cost without substantially sacrificing downstream performance. Existing data selection methods can remove redundancy but cannot repair imperfect supervision or supplement under-represented capabilities. We propose SANet, a fixed-budget data construction framework that alternates Multi-Criteria Selection and Bi-Level Augmentation. SANet refines imperfect instruction–response pairs, generates samples for under-represented capabilities, and re-evaluates augmented candidates through subsequent selection before constructing the final subset. Using only 6% of the original training data, SANet achieves an average score of 61.88 over the three predefined primary benchmarks (ScienceQA, MMBench, and TextVQA), outperforming 6% Random Selection by 1.83 points on this primary metric. On the additional diagnostic GQA benchmark, however, SANet performs worse than Random Selection, revealing a limitation in fine-grained compositional reasoning. The complete SANet pipeline requires approximately 35 h under the same hardware configuration, including a non-negligible one-time subset-construction cost of approximately 29 h and 6 h for downstream fine-tuning, compared with 67 h for full-data fine-tuning. Because the constructed subset can be reused across different backbones, the one-time construction overhead can be amortized when multiple downstream models are fine-tuned. The constructed subset can also be directly reused without backbone-specific reconstruction, achieving average scores of 69.03 on Qwen2-VL-7B and 87.56 on InternVL2.5-8B. These results indicate that SANet provides a data-efficient multimodal instruction-tuning pipeline with practical cross-model reusability. The reported multi-seed variation is based on three downstream fine-tuning runs with fixed constructed subsets and therefore reflects training-seed sensitivity rather than full-pipeline statistical uncertainty.

ElectronicsVol. 15(19)
Shanghai Jiao Tong University (CN), Chinese Academy of Sciences (CN), Beijing Academy of Artificial Intelligence (CN), Shenzhen Institutes of Advanced Technology (CN)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.