Consensus Ensembles for Zero-Shot Algorithm Classification with Large Language Models
Manual classification of algorithmic problems is costly and inconsistent, although it remains essential for curriculum sequencing, recommendation systems, and quality control on competitive programming platforms. To address this limitation, this study investigates whether consensus-based, zero-shot classification using large language models can replace manual annotation. Five architecturally diverse models were evaluated: Claude Sonnet v4.6 (CS4.6), DeepSeek v3.2 (DS3.2), Gemini v3 Flash (G3F), GPT v5.4 (GPT5.4), and Grok v4.1 Fast (GRK4.1). Each model independently classified 2911 LeetCode problems into 13 mutually exclusive algorithm categories to generate predicted labels and confidence vectors. These outputs served as the basis for systematically comparing eight consensus aggregation methods and five disagreement-resolution strategies. The five-model ensemble achieved at least 4/5 agreement on 85.3% of the problems, demonstrating substantial overall inter-rater reliability. While consensus weakened as problem difficulty increased, it was notably lowest for the Algorithmic Paradigms group because category boundaries depend on computational properties that are not directly visible in source code. Six aggregation methods converged within a narrow 90.0–90.1% average-agreement band. A Friedman omnibus test with Holm–Bonferroni-corrected pairwise comparisons showed that this band separates further into two statistically distinguishable trios differing by only 0.07–0.08 percentage points, illustrating that statistical significance at this sample size does not imply practical importance. An unsupervised adaptation of pairwise ranking from LLM-Blender reached 89.1%, whereas Borda Count underperformed at 83.1%. Among disagreement-resolution strategies, progressive ensemble pruning proved most effective by resolving 92.5% of cases. Calibration varied substantially across the panel, with GPT5.4 showing the largest calibration gap (0.124) and GRK4.1 showing the smallest (0.023). This variation suggests that confidence-weighted aggregation is effective only when applied selectively to well-calibrated models. These findings support consensus-based ensembles of these five models as a scalable alternative to manual algorithmic annotation and provide actionable guidance for designing robust aggregation pipelines in unsupervised settings.
Authors
- Haydar Tuna (ORCID: https://orcid.org/0000-0003-2388-653X)
- Sefa Akca (ORCID: https://orcid.org/0000-0002-8642-0537)
- Taymaz Akan (ORCID: https://orcid.org/0000-0003-4070-1058)
Institutions
- Osmaniye Korkut Ata University (TR)
- Louisiana State University Health Sciences Center Shreveport (US)
Publication Details
- Journal
- Electronics
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/electronics15194548
- Primary Topic
- Software Engineering Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00