Content-aware malicious website detection using multilingual BERT

Abstract Malicious website detection traditionally relies on URL analysis, but this is limited as attackers can create clean-looking links, abuse URL shorteners, or even hijack trusted websites. To address this gap, we shift the focus from analyzing technical artifacts to understanding intent, by proposing a content-aware approach that analyzes the actual text users see alongside the URL to capture malicious signals at a deeper level. We designed a preprocessing pipeline that converts raw HTML to concise Markdown format, reducing input length by over 96% while preserving essential content and structure. Using a large, multilingual dataset of 689,556 webpages, we fine-tuned and evaluated multiple BERT-based models under URL-only and URL+content settings, as well as traditional machine learning methods and hybrid ensembles. Results show that introducing content significantly improves performance, allowing even a simple TF-IDF + logistic regression model to surpass all URL-only transformers in accuracy. The best content-aware transformer, XLM-RoBERTa Base, achieved 99.01% accuracy, a 62.2% error reduction over its URL-only counterpart. Through hybrid ensembling, its accuracy climbs to 99.06% with a false positive rate of 0.70% and false negative rate of 1.20%, outperforming prior research across multiple benchmark datasets by 0.78 to 7.70 percentage points. Our findings demonstrate that reliable verdicts require appropriate context, and with the right inputs, even lightweight solutions can be robust across multiple languages and attack types.

Authors

Institutions

Publication Details

Journal
Discover Computing
Published
2026-10-03
DOI
https://doi.org/10.1007/s10791-026-10641-9
Primary Topic
Spam and Phishing Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Content-aware malicious website detection using multilingual BERT

Khuong Nguyen‐Vinh, Minh Anh Hoang, Cuong Nguyen Viet Le
Discover Computing
Spam and Phishing Detection
article

Content-aware malicious website detection using multilingual BERT

Khuong Nguyen‐Vinh, Minh Anh Hoang, Cuong Nguyen Viet Le
article en

Abstract

Abstract Malicious website detection traditionally relies on URL analysis, but this is limited as attackers can create clean-looking links, abuse URL shorteners, or even hijack trusted websites. To address this gap, we shift the focus from analyzing technical artifacts to understanding intent, by proposing a content-aware approach that analyzes the actual text users see alongside the URL to capture malicious signals at a deeper level. We designed a preprocessing pipeline that converts raw HTML to concise Markdown format, reducing input length by over 96% while preserving essential content and structure. Using a large, multilingual dataset of 689,556 webpages, we fine-tuned and evaluated multiple BERT-based models under URL-only and URL+content settings, as well as traditional machine learning methods and hybrid ensembles. Results show that introducing content significantly improves performance, allowing even a simple TF-IDF + logistic regression model to surpass all URL-only transformers in accuracy. The best content-aware transformer, XLM-RoBERTa Base, achieved 99.01% accuracy, a 62.2% error reduction over its URL-only counterpart. Through hybrid ensembling, its accuracy climbs to 99.06% with a false positive rate of 0.70% and false negative rate of 1.20%, outperforming prior research across multiple benchmark datasets by 0.78 to 7.70 percentage points. Our findings demonstrate that reliable verdicts require appropriate context, and with the right inputs, even lightweight solutions can be robust across multiple languages and attack types.

Discover ComputingVol. 29(1)
FPT University (VN), VSB - Technical University of Ostrava (CZ), RMIT Vietnam (VN)
Openalex Percentile: Top 5%
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Content-aware malicious website detection using multilingual BERT — Khuong Nguyen‐Vinh, Minh Anh Hoang, et al. · Discover Computing (2026) | TGRS Research Map | TGRS