Neuro-Symbolic Sentence Ranking for Malayalam Extractive Summarization with Conservative Dynamic MMR

Malayalam remains comparatively under-resourced for sentence-level summarization, while its agglutinative morphology, complex orthography, and news-writing conventions create additional modeling and deployment challenges. This paper presents a resource-conscious extractive summarization pipeline that ranks complete Malayalam source sentences instead of generating new text. The system evolved from an unsupervised LaBSE centroid baseline to a supervised sentence classifier and finally to a dual-path neuro-symbolic architecture. The semantic branch consumes a frozen multilingual sentence embedding; the symbolic branch encodes sentence position, normalized length, complex-word density, and numeral density. Their fused representation estimates sentence salience. A refactored Dynamic Maximal Marginal Relevance (D-MMR) selector normalizes relevance and cosine-similarity scales, bounds heuristic contributions, and protects classifier ranking through relevance-retention checks and fallback behavior. Five checkpoint and encoder variants were compared on a 50-article four-column consensus development set under oracle-length and production-length conditions. The Chotta Bheem checkpoint achieved the highest observed oracle-length Micro F1 of 0.6241 and Macro F1 of 0.6200, and the highest production-length Micro F1 of 0.5141. Chotta Bheem V2 remained qualitatively useful on broader frontend examples but did not improve aggregate sentence-index overlap. The results establish Chotta Bheem as the strongest current deployment checkpoint while also showing that the development set is not a substitute for a new, document-disjoint human benchmark.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23138411
Primary Topic
Advanced Text Analysis Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Neuro-Symbolic Sentence Ranking for Malayalam Extractive Summarization with Conservative Dynamic MMR

Adithya Kiran
Zenodo (CERN European Organization for Nuclear Research)
Advanced Text Analysis Techniques
preprint

Neuro-Symbolic Sentence Ranking for Malayalam Extractive Summarization with Conservative Dynamic MMR

Adithya Kiran
preprint en

Abstract

Malayalam remains comparatively under-resourced for sentence-level summarization, while its agglutinative morphology, complex orthography, and news-writing conventions create additional modeling and deployment challenges. This paper presents a resource-conscious extractive summarization pipeline that ranks complete Malayalam source sentences instead of generating new text. The system evolved from an unsupervised LaBSE centroid baseline to a supervised sentence classifier and finally to a dual-path neuro-symbolic architecture. The semantic branch consumes a frozen multilingual sentence embedding; the symbolic branch encodes sentence position, normalized length, complex-word density, and numeral density. Their fused representation estimates sentence salience. A refactored Dynamic Maximal Marginal Relevance (D-MMR) selector normalizes relevance and cosine-similarity scales, bounds heuristic contributions, and protects classifier ranking through relevance-retention checks and fallback behavior. Five checkpoint and encoder variants were compared on a 50-article four-column consensus development set under oracle-length and production-length conditions. The Chotta Bheem checkpoint achieved the highest observed oracle-length Micro F1 of 0.6241 and Macro F1 of 0.6200, and the highest production-length Micro F1 of 0.5141. Chotta Bheem V2 remained qualitatively useful on broader frontend examples but did not improve aggregate sentence-index overlap. The results establish Chotta Bheem as the strongest current deployment checkpoint while also showing that the development set is not a substitute for a new, document-disjoint human benchmark.

Zenodo (CERN European Organization for Nuclear Research)
TKM College of Engineering (IN)
Advanced Text Analysis Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.