Efficient Multimodal Instruction Tuning Through Iterative Data Selection and Augmentation Under Fixed Data Budgets
Efficient multimodal instruction tuning requires reducing both training data volume and computational cost without substantially sacrificing downstream performance. Existing data selection methods can remove redundancy but cannot repair imperfect supervision or supplement under-represented capabilities. We propose SANet, a fixed-budget data construction framework that alternates Multi-Criteria Selection and Bi-Level Augmentation. SANet refines imperfect instruction–response pairs, generates samples for under-represented capabilities, and re-evaluates augmented candidates through subsequent selection before constructing the final subset. Using only 6% of the original training data, SANet achieves an average score of 61.88 over the three predefined primary benchmarks (ScienceQA, MMBench, and TextVQA), outperforming 6% Random Selection by 1.83 points on this primary metric. On the additional diagnostic GQA benchmark, however, SANet performs worse than Random Selection, revealing a limitation in fine-grained compositional reasoning. The complete SANet pipeline requires approximately 35 h under the same hardware configuration, including a non-negligible one-time subset-construction cost of approximately 29 h and 6 h for downstream fine-tuning, compared with 67 h for full-data fine-tuning. Because the constructed subset can be reused across different backbones, the one-time construction overhead can be amortized when multiple downstream models are fine-tuned. The constructed subset can also be directly reused without backbone-specific reconstruction, achieving average scores of 69.03 on Qwen2-VL-7B and 87.56 on InternVL2.5-8B. These results indicate that SANet provides a data-efficient multimodal instruction-tuning pipeline with practical cross-model reusability. The reported multi-seed variation is based on three downstream fine-tuning runs with fixed constructed subsets and therefore reflects training-seed sensitivity rather than full-pipeline statistical uncertainty.
Authors
- Yi Wang (ORCID: https://orcid.org/0000-0001-8659-4724)
- H Y Lin (ORCID: https://orcid.org/0000-0001-8048-7522)
- Yali Wang
- Yinan He
Institutions
- Shanghai Jiao Tong University (CN)
- Chinese Academy of Sciences (CN)
- Beijing Academy of Artificial Intelligence (CN)
- Shenzhen Institutes of Advanced Technology (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/electronics15194461
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00