Vision-language sample scoring with dynamic coverage-centric strategies

Data pruning reduces image-classifier training cost by replacing a full training set with a coreset. Existing methods often require training a target or proxy model to estimate sample importance, which ties data selection to a particular architecture. In addition, common allocation rules do not explicitly change the selected score distribution with the retention budget. We study how to select a fixed-size subset before target-model training while enabling score reuse across architectures and budget-dependent allocation. For this image-classification data-pruning task, we use a frozen Contrastive Language–Image Pre-training (CLIP) encoder as the artificial-intelligence component of a two-stage framework. CLIPSelector computes a two-prompt image-label alignment probability for each training image without optimizing a target or proxy classifier. Dynamic Coverage-centric Coreset Selection (DCCS) partitions the ranked scores into equal-count strata, allocates the initial budget to higher-scoring strata, and samples the remaining budget from a merged lower-scoring region using importance-weighted and uniform probabilities. Experiments on three image-classification benchmarks show that CLIPSelector-DCCS achieves the best average rank among the evaluated methods, generally improves on coverage-centric allocation under matched scoring functions, supports score reuse across target architectures and retention budgets, and reduces downstream training time. Separating score generation from coreset construction enables multiple fixed-budget coresets to be constructed without retraining a scoring model.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-09-14
DOI
https://doi.org/10.1016/j.engappai.2026.116201
Primary Topic
Domain Adaptation and Few-Shot Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Vision-language sample scoring with dynamic coverage-centric strategies

Mukesh Prasad, Xing Zi, Yunxiao Shi, Prayag Tiwari et al.
Engineering Applications of Artificial Intelligence
Domain Adaptation and Few-Shot Learning
article

Vision-language sample scoring with dynamic coverage-centric strategies

Mukesh Prasad, Xing Zi, Yunxiao Shi, Prayag Tiwari, Taoyuan Zhu, Xian Tao, Min Xu, Jun Li
article en

Abstract

Data pruning reduces image-classifier training cost by replacing a full training set with a coreset. Existing methods often require training a target or proxy model to estimate sample importance, which ties data selection to a particular architecture. In addition, common allocation rules do not explicitly change the selected score distribution with the retention budget. We study how to select a fixed-size subset before target-model training while enabling score reuse across architectures and budget-dependent allocation. For this image-classification data-pruning task, we use a frozen Contrastive Language–Image Pre-training (CLIP) encoder as the artificial-intelligence component of a two-stage framework. CLIPSelector computes a two-prompt image-label alignment probability for each training image without optimizing a target or proxy classifier. Dynamic Coverage-centric Coreset Selection (DCCS) partitions the ranked scores into equal-count strata, allocates the initial budget to higher-scoring strata, and samples the remaining budget from a merged lower-scoring region using importance-weighted and uniform probabilities. Experiments on three image-classification benchmarks show that CLIPSelector-DCCS achieves the best average rank among the evaluated methods, generally improves on coverage-centric allocation under matched scoring functions, supports score reuse across target architectures and retention budgets, and reduces downstream training time. Separating score generation from coreset construction enables multiple fixed-budget coresets to be constructed without retraining a scoring model.

Engineering Applications of Artificial IntelligenceVol. 183
University of Technology Sydney (AU), Chinese Academy of Sciences (CN), Shandong Institute of Automation (CN), Halmstad University (SE)
Openalex Percentile: Top 8%
Domain Adaptation and Few-Shot Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.