Hybrid Stacking Approach for Biomedical Full-Text Classification: Combining ILP with Refinement Operators and Propositional Learners
This study addresses the problem of automatic classification of MEDLINE full-text biomedical documents and investigates whether relational learning can enhance performance within a stacking-based text classification framework. A major problem in using Inductive Logic Programming systems is their limited scalability in the presence of large search spaces and with many examples. The research addresses (i) the incorporation of an Inductive Logic Programming component into a traditional machine learning pipeline, namely WEKA platform, and (ii) a new methodology to enable a significant reduction in the redundancy of the hypothesis space in Inductive Logic Programming systems (ILP). This reduction enabled the use of Inductive Logic Programming in the very large and complex full-text classification problems. To accomplish the first objective, the Aleph system—an Inductive Logic Programming system—was integrated into the WEKA platform, and for that, we used a stacking architecture in which propositional learners and Inductive Logic Programming were applied to different document sections, with their outputs combined by a meta-learner. Experiments were conducted on an extended OHSUMED corpus comprising MEDLINE full-text documents mapped to MeSH disease categories. The achieved results indicate the feasibility of using propositional learners with relational learners, taking advantage of an integrated platform. As far as the available literature indicates, a hybrid solution that combines Inductive Logic Programming and propositional learners within a multi-view stacking architecture for MEDLINE full-text classification constitutes a novel approach. The results obtained by the proposed reduction in the redundancy of the hypothesis space were very promising, leading to substantial improvements for several disease classes, with some configurations achieving almost perfect agreement regarding the kappa statistics metric.
Authors
- Célia Talma Gonçalves (ORCID: https://orcid.org/0000-0002-3861-0854)
- Eva Iglesias (ORCID: https://orcid.org/0000-0002-7172-6947)
- Adrián Seara Vieira (ORCID: https://orcid.org/0000-0002-0099-3546)
- Rui Camacho (ORCID: https://orcid.org/0000-0003-0940-3554)
- L. Borrajo (ORCID: https://orcid.org/0000-0002-6089-6166)
- Carlos Gonçalves (ORCID: https://orcid.org/0000-0001-8107-8747)
Institutions
- Instituto Superior de Contabilidade e Administracao do Porto (PT)
- Universidade do Porto (PT)
- INESC TEC (PT)
- Universidade de Vigo (ES)
- Polytechnic Institute of Porto (PT)
Publication Details
- Journal
- Information
- Published
- 2026-09-09
- DOI
- https://doi.org/10.3390/info17090872
- Primary Topic
- Biomedical Text Mining and Ontologies
- Type
- article
- Field-Weighted Citation Impact
- 0.00