Complementary-Label Learning: A Survey of Methods, an Open-Source Toolkit, and Fair Benchmarking Results

Complementary-label learning (CLL) is a weakly supervised learning paradigm designed for multiclass classification tasks. In CLL, algorithms are provided with complementary labels, which indicate the classes an instance does not belong to, rather than specifying its true class. With the growing research interest in CLL and the unique challenges it presents, particularly in handling incomplete and indirect supervision, a comprehensive and systematic overview of existing methodologies, results, and research directions has become increasingly essential. In this survey, we provide the first systematic synthesis of recent advances in CLL and offer three major contributions. First, we establish a structured taxonomy that clarifies foundational concepts, distinctive algorithmic strategies, key contributions, and their inherent limitations. Second, to support reproducibility and accelerate future research, we introduce a publicly available open-source CLL toolkit that unifies key algorithm implementations and offers ready-to-use benchmarking pipelines. Third, we conduct an extensive benchmarking study covering a wide spectrum of CLL algorithms, using standardized evaluation protocols across both synthetic, human-annotated, and VLM-annotated complementary-label datasets. This provides the most complete performance comparison currently available. Finally, we discuss current open research challenges and identify promising avenues for future investigation, emphasizing improvements to the practical applicability and robustness of CLL in real-world scenarios. Overall, this survey provides an up-to-date, in-depth resource on the evolving CLL landscape.

Authors

Institutions

Publication Details

Journal
ACM Computing Surveys
Published
2026-10-03
DOI
https://doi.org/10.1145/3847508
Primary Topic
Text and Document Classification Technologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Complementary-Label Learning: A Survey of Methods, an Open-Source Toolkit, and Fair Benchmarking Results

Masashi Sugiyama, Nai-Xuan Ye, Gang Niu, Tan-Ha Mai et al.
ACM Computing Surveys
Text and Document Classification Technologies
article

Complementary-Label Learning: A Survey of Methods, an Open-Source Toolkit, and Fair Benchmarking Results

Masashi Sugiyama, Nai-Xuan Ye, Gang Niu, Tan-Ha Mai, Hsuan-Tien Lin, Takashi Ishida, Wei Wang
article en

Abstract

Complementary-label learning (CLL) is a weakly supervised learning paradigm designed for multiclass classification tasks. In CLL, algorithms are provided with complementary labels, which indicate the classes an instance does not belong to, rather than specifying its true class. With the growing research interest in CLL and the unique challenges it presents, particularly in handling incomplete and indirect supervision, a comprehensive and systematic overview of existing methodologies, results, and research directions has become increasingly essential. In this survey, we provide the first systematic synthesis of recent advances in CLL and offer three major contributions. First, we establish a structured taxonomy that clarifies foundational concepts, distinctive algorithmic strategies, key contributions, and their inherent limitations. Second, to support reproducibility and accelerate future research, we introduce a publicly available open-source CLL toolkit that unifies key algorithm implementations and offers ready-to-use benchmarking pipelines. Third, we conduct an extensive benchmarking study covering a wide spectrum of CLL algorithms, using standardized evaluation protocols across both synthetic, human-annotated, and VLM-annotated complementary-label datasets. This provides the most complete performance comparison currently available. Finally, we discuss current open research challenges and identify promising avenues for future investigation, emphasizing improvements to the practical applicability and robustness of CLL in real-world scenarios. Overall, this survey provides an up-to-date, in-depth resource on the evolving CLL landscape.

ACM Computing Surveys
National Taiwan University (TW), RIKEN (JP), RIKEN Center for Advanced Intelligence Project (JP), The University of Tokyo (JP)
Openalex Percentile: Top 9%
Text and Document Classification Technologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.