Time series modeling and anomaly removal for Wikipedia web traffic prediction

Abstract Accurate forecasting of web traffic is essential for a web service provider who wants to manage infrastructure resources and decision-making processes. Internet users rely heavily on Wikipedia articles for information, and over the past few years, the usage of Wikipedia has increased significantly. Several forecasting and prediction models for specific web page traffic have been published previously that include statistical and deep learning methods. The web traffic data modeling of Wikipedia pages is challenging due to high dimensionality, seasonality, anomalies, and missing values. These anomalies, if left undetected, can adversely affect the overall performance of time series models. In this research, a comprehensive methodology for web traffic prediction is proposed that integrates robust preprocessing, particularly anomaly detection and removal using Isolation Forest, with a comparative analysis of eight time series forecasting models, including ARMA, ARIMA, Auto ARIMA, Exponential Smoothing, Prophet, LSTM, BiLSTM, and MA. The dataset for this research is a Kaggle competition dataset, and Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), and Mean Absolute Error (MAE) are utilized as the evaluation metrics to measure the model’s performance. The findings indicate that the Moving Average model performs better for this dataset when evaluated using RMSE, MAPE, and MAE value (percentage deviation) as compared to other models. A comparison of eight time series model performances is presented based on the Wikipedia dataset, which comprises 145,000 articles.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-06
DOI
https://doi.org/10.1038/s41598-026-73011-x
Primary Topic
Forecasting Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Time series modeling and anomaly removal for Wikipedia web traffic prediction

Sandhya Avasthi, Inung Wijayanto, Suman Lata Tripathi, Thein Kyaw LWIN
Scientific Reports
Forecasting Techniques and Applications
article

Time series modeling and anomaly removal for Wikipedia web traffic prediction

Sandhya Avasthi, Inung Wijayanto, Suman Lata Tripathi, Thein Kyaw LWIN
article en

Abstract

Abstract Accurate forecasting of web traffic is essential for a web service provider who wants to manage infrastructure resources and decision-making processes. Internet users rely heavily on Wikipedia articles for information, and over the past few years, the usage of Wikipedia has increased significantly. Several forecasting and prediction models for specific web page traffic have been published previously that include statistical and deep learning methods. The web traffic data modeling of Wikipedia pages is challenging due to high dimensionality, seasonality, anomalies, and missing values. These anomalies, if left undetected, can adversely affect the overall performance of time series models. In this research, a comprehensive methodology for web traffic prediction is proposed that integrates robust preprocessing, particularly anomaly detection and removal using Isolation Forest, with a comparative analysis of eight time series forecasting models, including ARMA, ARIMA, Auto ARIMA, Exponential Smoothing, Prophet, LSTM, BiLSTM, and MA. The dataset for this research is a Kaggle competition dataset, and Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), and Mean Absolute Error (MAE) are utilized as the evaluation metrics to measure the model’s performance. The findings indicate that the Moving Average model performs better for this dataset when evaluated using RMSE, MAPE, and MAE value (percentage deviation) as compared to other models. A comparison of eight time series model performances is presented based on the Wikipedia dataset, which comprises 145,000 articles.

Scientific Reports
Symbiosis International University (IN), Batangas State University (PH), ABES Engineering College, Telkom University (ID)
Openalex Percentile: Top 9%
Forecasting Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Time series modeling and anomaly removal for Wikipedia web traffic prediction — Sandhya Avasthi, Inung Wijayanto, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS