Prompt Escalation for Lightweight Large Language Models: An Empirical Evaluation of Cost–Performance Trade-Offs
Prompt escalation can increase the resource requirements of lightweight large language models (LLMs) without improving predictive performance. We evaluated four instruction-tuned 2–4B models on Grade School Math 8K (GSM8K), CommonsenseQA (CSQA), Recognizing Textual Entailment (RTE), and the binary Stanford Sentiment Treebank (SST-2). Zero-shot (ZS), few-shot (FS), chain-of-thought (CoT), and few-shot CoT (FS+CoT) were compared across 64 conditions using predictive performance, tokens, latency, and stochastic response consistency. Demonstrations came from training splits; FS+CoT used worked rationales with automatic screening and a partial manual audit. Latency was measured separately with synchronization, warm-up exclusion, and counterbalanced prompt order. After Holm adjustment, 12.5% of predictive-performance contrasts were significant, and ZS was significantly outperformed in 1 of 48 contrasts, compared with significant differences in 100% of total-token and 75.0% of latency contrasts. Strategy rankings varied by model and task. A single auxiliary 7B model showed no significant accuracy gain over ZS but did not establish a general scale effect. Scenario-weight sensitivity frequently favored ZS, with model–task exceptions. These results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings. The rankings and guidelines have not been validated for summarization, code generation, or multi-turn dialogue.
Authors
- Bonggyun Ko (ORCID: https://orcid.org/0000-0002-1544-6377)
- Seyoung Kim (ORCID: https://orcid.org/0000-0003-2197-4755)
Institutions
- Chonnam National University (KR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/app16189190
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00