Detection Of AI: Generated Phishing Emails Using Cross Model Generalization

Large language models (LLMs) have changed the practical characteristics of phishing email.A message that once required deliberate manual drafting can now be produced in seconds, adapted to a target role, and rewritten repeatedly while maintaining fluent grammar and a credible business tone.This development weakens the assumption behind many older filters: that phishing is recognizable because it is poorly written, repetitive, or linguistically unusual.The central question of this paper is therefore not whether artificial intelligence can detect phishing, but whether current AI-assisted detection approaches remain reliable when the generator, wording, context, and delivery conditions change.This paper presents a structured gap analysis of recent approaches to detecting AI-and LLM-generated phishing emails.The review focuses on four recurring dimensions: detection capability, cross-model generalization, operational latency, and interpretability.Recent evidence shows that stylometric systems can perform strongly when the evaluation distribution is close to the training distribution, while broader reviews identify over-reliance on manually assembled datasets and disproportionate attention to GPT-family models.A 2026 cross-model study further demonstrates that a detector trained on one generator can lose substantial transfer performance when evaluated on another, although threshold recalibration and multi-generator training can reduce the observed gap [1], [2], [3].The analysis argues that the strongest and most defensible research gap is cross-model robustness under distribution shift.Latency and explainability remain important engineering constraints, but generalization is the issue that most directly challenges the scientific validity of a detector if its benchmark assumes a single or narrow generator family.The paper therefore treats generalization as the primary gap and positions latency, false-positive control, privacy, and explanation quality as secondary deployment requirements.Rather than claiming completed experimental results, the paper

Authors

Publication Details

Journal
International Journal of Innovative Research in Technology
Published
2026-09-15
DOI
https://doi.org/10.64643/ijirt.208482-459
Primary Topic
Spam and Phishing Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Detection Of AI: Generated Phishing Emails Using Cross Model Generalization

Guna Dhondwad, Shamsher Singh Hada, Shruti Dhotre, Alokananda Ghosh
International Journal of Innovative Research in Technology
Spam and Phishing Detection
article

Detection Of AI: Generated Phishing Emails Using Cross Model Generalization

Guna Dhondwad, Shamsher Singh Hada, Shruti Dhotre, Alokananda Ghosh
article en

Abstract

Large language models (LLMs) have changed the practical characteristics of phishing email.A message that once required deliberate manual drafting can now be produced in seconds, adapted to a target role, and rewritten repeatedly while maintaining fluent grammar and a credible business tone.This development weakens the assumption behind many older filters: that phishing is recognizable because it is poorly written, repetitive, or linguistically unusual.The central question of this paper is therefore not whether artificial intelligence can detect phishing, but whether current AI-assisted detection approaches remain reliable when the generator, wording, context, and delivery conditions change.This paper presents a structured gap analysis of recent approaches to detecting AI-and LLM-generated phishing emails.The review focuses on four recurring dimensions: detection capability, cross-model generalization, operational latency, and interpretability.Recent evidence shows that stylometric systems can perform strongly when the evaluation distribution is close to the training distribution, while broader reviews identify over-reliance on manually assembled datasets and disproportionate attention to GPT-family models.A 2026 cross-model study further demonstrates that a detector trained on one generator can lose substantial transfer performance when evaluated on another, although threshold recalibration and multi-generator training can reduce the observed gap [1], [2], [3].The analysis argues that the strongest and most defensible research gap is cross-model robustness under distribution shift.Latency and explainability remain important engineering constraints, but generalization is the issue that most directly challenges the scientific validity of a detector if its benchmark assumes a single or narrow generator family.The paper therefore treats generalization as the primary gap and positions latency, false-positive control, privacy, and explanation quality as secondary deployment requirements.Rather than claiming completed experimental results, the paper

International Journal of Innovative Research in TechnologyVol. 13(5)
Quality Education
Openalex Percentile: Top 4%
Spam and Phishing Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.