Consensus Ensembles for Zero-Shot Algorithm Classification with Large Language Models

Manual classification of algorithmic problems is costly and inconsistent, although it remains essential for curriculum sequencing, recommendation systems, and quality control on competitive programming platforms. To address this limitation, this study investigates whether consensus-based, zero-shot classification using large language models can replace manual annotation. Five architecturally diverse models were evaluated: Claude Sonnet v4.6 (CS4.6), DeepSeek v3.2 (DS3.2), Gemini v3 Flash (G3F), GPT v5.4 (GPT5.4), and Grok v4.1 Fast (GRK4.1). Each model independently classified 2911 LeetCode problems into 13 mutually exclusive algorithm categories to generate predicted labels and confidence vectors. These outputs served as the basis for systematically comparing eight consensus aggregation methods and five disagreement-resolution strategies. The five-model ensemble achieved at least 4/5 agreement on 85.3% of the problems, demonstrating substantial overall inter-rater reliability. While consensus weakened as problem difficulty increased, it was notably lowest for the Algorithmic Paradigms group because category boundaries depend on computational properties that are not directly visible in source code. Six aggregation methods converged within a narrow 90.0–90.1% average-agreement band. A Friedman omnibus test with Holm–Bonferroni-corrected pairwise comparisons showed that this band separates further into two statistically distinguishable trios differing by only 0.07–0.08 percentage points, illustrating that statistical significance at this sample size does not imply practical importance. An unsupervised adaptation of pairwise ranking from LLM-Blender reached 89.1%, whereas Borda Count underperformed at 83.1%. Among disagreement-resolution strategies, progressive ensemble pruning proved most effective by resolving 92.5% of cases. Calibration varied substantially across the panel, with GPT5.4 showing the largest calibration gap (0.124) and GRK4.1 showing the smallest (0.023). This variation suggests that confidence-weighted aggregation is effective only when applied selectively to well-calibrated models. These findings support consensus-based ensembles of these five models as a scalable alternative to manual algorithmic annotation and provide actionable guidance for designing robust aggregation pipelines in unsupervised settings.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-06
DOI
https://doi.org/10.3390/electronics15194548
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Consensus Ensembles for Zero-Shot Algorithm Classification with Large Language Models

Haydar Tuna, Sefa Akca, Taymaz Akan
Electronics
Software Engineering Research
article

Consensus Ensembles for Zero-Shot Algorithm Classification with Large Language Models

Haydar Tuna, Sefa Akca, Taymaz Akan
article en

Abstract

Manual classification of algorithmic problems is costly and inconsistent, although it remains essential for curriculum sequencing, recommendation systems, and quality control on competitive programming platforms. To address this limitation, this study investigates whether consensus-based, zero-shot classification using large language models can replace manual annotation. Five architecturally diverse models were evaluated: Claude Sonnet v4.6 (CS4.6), DeepSeek v3.2 (DS3.2), Gemini v3 Flash (G3F), GPT v5.4 (GPT5.4), and Grok v4.1 Fast (GRK4.1). Each model independently classified 2911 LeetCode problems into 13 mutually exclusive algorithm categories to generate predicted labels and confidence vectors. These outputs served as the basis for systematically comparing eight consensus aggregation methods and five disagreement-resolution strategies. The five-model ensemble achieved at least 4/5 agreement on 85.3% of the problems, demonstrating substantial overall inter-rater reliability. While consensus weakened as problem difficulty increased, it was notably lowest for the Algorithmic Paradigms group because category boundaries depend on computational properties that are not directly visible in source code. Six aggregation methods converged within a narrow 90.0–90.1% average-agreement band. A Friedman omnibus test with Holm–Bonferroni-corrected pairwise comparisons showed that this band separates further into two statistically distinguishable trios differing by only 0.07–0.08 percentage points, illustrating that statistical significance at this sample size does not imply practical importance. An unsupervised adaptation of pairwise ranking from LLM-Blender reached 89.1%, whereas Borda Count underperformed at 83.1%. Among disagreement-resolution strategies, progressive ensemble pruning proved most effective by resolving 92.5% of cases. Calibration varied substantially across the panel, with GPT5.4 showing the largest calibration gap (0.124) and GRK4.1 showing the smallest (0.023). This variation suggests that confidence-weighted aggregation is effective only when applied selectively to well-calibrated models. These findings support consensus-based ensembles of these five models as a scalable alternative to manual algorithmic annotation and provide actionable guidance for designing robust aggregation pipelines in unsupervised settings.

ElectronicsVol. 15(19)
Osmaniye Korkut Ata University (TR), Louisiana State University Health Sciences Center Shreveport (US)
Openalex Percentile: Top 5%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.