Modeling the Cognitive Effects of Social Media Algorithms With Textual and Behavioral Data: A Deep Hybrid DistilBERT , BiLSTM , and CatBoost Analysis

ABSTRACT The basic concept of the algorithmic ranking systems in the social media setting is to influence what people watch, read, and eventually think. Focusing on engagement, such algorithms unwillingly intensify selective exposure and modeling behavioral proxies of user attention and selective exposure, but the process by which content curation is converted into different attention has not been empirically studied. In this research, a hybrid DistilBERT‐BiLSTM‐CatBoost model is suggested to quantitatively examine the joint effect of semantic features and behavioral history on user attention and cognition in an algorithmic feed. The framework works at the slate level using the Microsoft news dataset (MIND) that includes impressions level exposure logs and rich textual metadata, but relies on three complementary processes: (1) DistilBERT encodes the semantic content of news titles and abstracts; (2) BiLSTM takes into consideration the sequential user preferences and short term behavioral dynamics; (3) CatBoost is a bias aware learning with the capability to rank semantic content by means of textual, sequential, and contextual covariates, such as position, popularity, rec Counterfactual evaluation techniques (propensity weighted) (IPS weighted NDCG) are used to identify the real relevance and exposure effects. The experimental findings indicate that the proposed model is more superior as compared to five states of the art baselines (XGBoost, LightGBM, Random Forest, Decision Tree, and AdaBoost) with an F1 score of 0.956 and NDCG at 5 of 0.872. The effectiveness of the DistilBERT embedding in generalization is established by comparing across five deep architectures (LSTM, BiLSTM, GRU, CNN BiLSTM, and Transformer Encoder). SHAP attributed feature identification shows position and popularity to be the strongest predictors of attention, then semantic similarity and recency, the latter continuing to have a strong effect on user cognition despite bias correction. The results overall give a replicable, empirically‐driven model of bias conscious ranking analysis in personalized feeds, which provides insights on content algorithms in a scalable way in both maintaining engagement and contributing to the passive process of rationalizing modeling behavioral proxies of user attention and selective exposure.

Authors

Institutions

Publication Details

Journal
Computational Intelligence
Published
2026-09-25
DOI
https://doi.org/10.1111/coin.70300
Primary Topic
Sentiment Analysis and Opinion Mining
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Modeling the Cognitive Effects of Social Media Algorithms With Textual and Behavioral Data: A Deep Hybrid DistilBERT , BiLSTM , and CatBoost Analysis

Chao Hua Yan
Computational Intelligence
Sentiment Analysis and Opinion Mining
article

Modeling the Cognitive Effects of Social Media Algorithms With Textual and Behavioral Data: A Deep Hybrid DistilBERT , BiLSTM , and CatBoost Analysis

Chao Hua Yan
article en

Abstract

ABSTRACT The basic concept of the algorithmic ranking systems in the social media setting is to influence what people watch, read, and eventually think. Focusing on engagement, such algorithms unwillingly intensify selective exposure and modeling behavioral proxies of user attention and selective exposure, but the process by which content curation is converted into different attention has not been empirically studied. In this research, a hybrid DistilBERT‐BiLSTM‐CatBoost model is suggested to quantitatively examine the joint effect of semantic features and behavioral history on user attention and cognition in an algorithmic feed. The framework works at the slate level using the Microsoft news dataset (MIND) that includes impressions level exposure logs and rich textual metadata, but relies on three complementary processes: (1) DistilBERT encodes the semantic content of news titles and abstracts; (2) BiLSTM takes into consideration the sequential user preferences and short term behavioral dynamics; (3) CatBoost is a bias aware learning with the capability to rank semantic content by means of textual, sequential, and contextual covariates, such as position, popularity, rec Counterfactual evaluation techniques (propensity weighted) (IPS weighted NDCG) are used to identify the real relevance and exposure effects. The experimental findings indicate that the proposed model is more superior as compared to five states of the art baselines (XGBoost, LightGBM, Random Forest, Decision Tree, and AdaBoost) with an F1 score of 0.956 and NDCG at 5 of 0.872. The effectiveness of the DistilBERT embedding in generalization is established by comparing across five deep architectures (LSTM, BiLSTM, GRU, CNN BiLSTM, and Transformer Encoder). SHAP attributed feature identification shows position and popularity to be the strongest predictors of attention, then semantic similarity and recency, the latter continuing to have a strong effect on user cognition despite bias correction. The results overall give a replicable, empirically‐driven model of bias conscious ranking analysis in personalized feeds, which provides insights on content algorithms in a scalable way in both maintaining engagement and contributing to the passive process of rationalizing modeling behavioral proxies of user attention and selective exposure.

Computational IntelligenceVol. 42(5)
Nanyang Institute of Technology (CN)
Quality Education
Openalex Percentile: Top 9%
Sentiment Analysis and Opinion Mining
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.