Knowledge Driven Event Extraction from Low Resource Language Content Using Strategic Few-Shot Learning

Timely extracting structured information from social media during disaster events is critical for effective response and coordination. This work presents a comprehensive pipeline for sub-event information extraction (IE) from Hindi-language tweets and news articles, covering eleven key categories such as casualties, infrastructure damage, relief operations and others. We introduce CrisisIE-Hindi, a benchmark dataset of around 9.6k annotated posts spanning cyclone, earthquake, flood and other disaster events. To assess the robustness of disaster IE models under low-resource and multilingual settings, we evaluate seven open-source and proprietary language models of varying sizes. Our analysis reveals that smaller models struggle to follow structured output instructions and generalize across unseen scenarios, particularly those with limited exposure to Hindi or domain-specific data. To mitigate these challenges, we propose a strategic few-shot prompting framework that leverages a frozen data bank for dynamic demonstration selection. This approach significantly improves model performance across all eleven categories, even on out-of-domain events. To show the efficacy of our proposed approach, we also showed the performance on three other low-resource language datasets- L3-Cube-MahaNER, B-NER, and HiNER. Notably, our method enables small and mid-sized language models to generalize effectively in instruction-following tasks for disaster IE. Our findings highlight the potential of prompt-based adaptation for low-resource languages and domains. We release our dataset to foster future Hindi disaster response systems research.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Asian and Low-Resource Language Information Processing
Published
2026-09-24
DOI
https://doi.org/10.1145/3849475
Primary Topic
Public Relations and Crisis Communication
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowledge Driven Event Extraction from Low Resource Language Content Using Strategic Few-Shot Learning

Suman Kundu, Prasenjeet Roy, Lipika Dey
ACM Transactions on Asian and Low-Resource Language Information Processing
Public Relations and Crisis Communication
article

Knowledge Driven Event Extraction from Low Resource Language Content Using Strategic Few-Shot Learning

Suman Kundu, Prasenjeet Roy, Lipika Dey
article en

Abstract

Timely extracting structured information from social media during disaster events is critical for effective response and coordination. This work presents a comprehensive pipeline for sub-event information extraction (IE) from Hindi-language tweets and news articles, covering eleven key categories such as casualties, infrastructure damage, relief operations and others. We introduce CrisisIE-Hindi, a benchmark dataset of around 9.6k annotated posts spanning cyclone, earthquake, flood and other disaster events. To assess the robustness of disaster IE models under low-resource and multilingual settings, we evaluate seven open-source and proprietary language models of varying sizes. Our analysis reveals that smaller models struggle to follow structured output instructions and generalize across unseen scenarios, particularly those with limited exposure to Hindi or domain-specific data. To mitigate these challenges, we propose a strategic few-shot prompting framework that leverages a frozen data bank for dynamic demonstration selection. This approach significantly improves model performance across all eleven categories, even on out-of-domain events. To show the efficacy of our proposed approach, we also showed the performance on three other low-resource language datasets- L3-Cube-MahaNER, B-NER, and HiNER. Notably, our method enables small and mid-sized language models to generalize effectively in instruction-following tasks for disaster IE. Our findings highlight the potential of prompt-based adaptation for low-resource languages and domains. We release our dataset to foster future Hindi disaster response systems research.

ACM Transactions on Asian and Low-Resource Language Information Processing
Indian Institute of Technology Jodhpur (IN), Ashoka University (IN)
Climate action
Openalex Percentile: Top 4%
Public Relations and Crisis Communication
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Knowledge Driven Event Extraction from Low Resource Language Content Using Strategic Few-Shot Learning — Suman Kundu, Prasenjeet Roy, et al. · ACM Transactions on Asian and Low-Resource Language Information Processing (2026) | TGRS Research Map | TGRS