Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor selection framework

Dynamic Searchable Symmetric Encryption (DSSE) enables keyword search over encrypted data without revealing plaintext to the server. Arabic morphological richness — where a single root generates dozens of surface forms — creates substantial challenges for encrypted search: unprocessed vocabularies inflate encrypted-index size and transmission cost, and limit retrieval recall by failing to match morphological variants. This paper presents the first empirically grounded preprocessor selection framework for Arabic DSSE, derived from a systematic evaluation of four Arabic preprocessing strategies — the Khoja stemmer, ISRI stemmer, Lucene Arabic Analyzer, and Farasa segmenter — against a normalization-only baseline within the Incidence Matrix DSSE (IM-DSSE) scheme on the Khaleej Arabic news corpus. We evaluate vocabulary size, search latency, search quality, and retrieval breadth on an annotated 1,400-document corpus using 56 benchmark queries, and assess scalability on cloud infrastructure across corpus sizes up to 45,500 documents. Stemming reduces vocabulary by up to 83% relative to the normalization-only baseline, proportionally reducing encrypted-index size and transmission cost. We find that per-query search latency is driven primarily by result-set size rather than vocabulary size: the normalization-only baseline is fastest per query because it matches the fewest documents, while preprocessors that broaden retrieval incur higher latency. At scale, all five configurations — including the normalization-only baseline — operate successfully to 45,500 documents, and the result-set-size effect on latency persists. In terms of search quality, light stemming and morphological segmentation achieve the best trade-off between retrieval precision and recall, while root-based stemmers sacrifice precision without commensurate recall gains. Based on these findings, the framework provides actionable, evidence-based guidelines mapping deployment requirements — memory scalability, search latency, search quality, and retrieval breadth — to concrete preprocessor choices for Arabic DSSE system designers.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-21
DOI
https://doi.org/10.1371/journal.pone.0351281
Primary Topic
Cryptography and Data Security
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor selection framework

Kholoud Al-Saleh
PLoS ONE
Cryptography and Data Security
article

Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor selection framework

Kholoud Al-Saleh
article en

Abstract

Dynamic Searchable Symmetric Encryption (DSSE) enables keyword search over encrypted data without revealing plaintext to the server. Arabic morphological richness — where a single root generates dozens of surface forms — creates substantial challenges for encrypted search: unprocessed vocabularies inflate encrypted-index size and transmission cost, and limit retrieval recall by failing to match morphological variants. This paper presents the first empirically grounded preprocessor selection framework for Arabic DSSE, derived from a systematic evaluation of four Arabic preprocessing strategies — the Khoja stemmer, ISRI stemmer, Lucene Arabic Analyzer, and Farasa segmenter — against a normalization-only baseline within the Incidence Matrix DSSE (IM-DSSE) scheme on the Khaleej Arabic news corpus. We evaluate vocabulary size, search latency, search quality, and retrieval breadth on an annotated 1,400-document corpus using 56 benchmark queries, and assess scalability on cloud infrastructure across corpus sizes up to 45,500 documents. Stemming reduces vocabulary by up to 83% relative to the normalization-only baseline, proportionally reducing encrypted-index size and transmission cost. We find that per-query search latency is driven primarily by result-set size rather than vocabulary size: the normalization-only baseline is fastest per query because it matches the fewest documents, while preprocessors that broaden retrieval incur higher latency. At scale, all five configurations — including the normalization-only baseline — operate successfully to 45,500 documents, and the result-set-size effect on latency persists. In terms of search quality, light stemming and morphological segmentation achieve the best trade-off between retrieval precision and recall, while root-based stemmers sacrifice precision without commensurate recall gains. Based on these findings, the framework provides actionable, evidence-based guidelines mapping deployment requirements — memory scalability, search latency, search quality, and retrieval breadth — to concrete preprocessor choices for Arabic DSSE system designers.

PLoS ONEVol. 21(9)
King Saud University (SA)
Industry, innovation and infrastructure
Openalex Percentile: Top 8%
Cryptography and Data Security
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor selection framework — Kholoud Al-Saleh · PLoS ONE (2026) | TGRS Research Map | TGRS