TDPO: two-stage differentiated prompt optimization for image classification
Prompt engineering has emerged as an effective paradigm for adapting vision-language models (VLMs) to image classification. Although large language models (LLMs) could generate rich visual descriptions, their hallucination often produces inaccurate or weakly discriminative class-specific prompts. Moreover, existing prompt optimization methods uniformly optimize all classes, overlooking inter-class heterogeneity. To address these limitations, we propose a two-stage differentiated prompt optimization (TDPO) that refines prompts from task-specific templates to class-specific descriptions. In the first stage, TDPO optimizes templates via roulette-wheel selection to balance efficiency and diversity. In the second stage, salient classes are sampled via a class group sampling strategy, and prompts are optimized using hybrid breeding optimization algorithm (HBO). This cooperative algorithm assigns class-specific prompts to three evolutionary lines based on inter-class heterogeneity, enabling the efficient discovery of discriminative class-specific prompts while preserving stable prompts for well-performing classes. In the setting of challenging one-shot image classification, extensive experiments on seven image classification datasets demonstrate that the proposed TDPO effectively improves classification accuracy, particularly on the Flowers102 and DTD datasets with gains of 2.9% and 1.7% over the previous state of the art. This work provides a promising direction for prompt optimization in vision-language applications.
Authors
- Mengqing Mei
- Songsong Zhang (ORCID: https://orcid.org/0009-0008-3203-878X)
- Zhiwei Ye
- Jia Guo
- Qubo Xie (ORCID: https://orcid.org/0009-0001-2232-5055)
- Yunjie Zeng
- Qiyi He
Institutions
- Wuhan Donghu University (CN)
- Nagoya University (JP)
- Hubei University of Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-26
- DOI
- https://doi.org/10.1038/s41598-026-72568-x
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00