Modeling the Cognitive Effects of Social Media Algorithms With Textual and Behavioral Data: A Deep Hybrid DistilBERT , BiLSTM , and CatBoost Analysis
ABSTRACT The basic concept of the algorithmic ranking systems in the social media setting is to influence what people watch, read, and eventually think. Focusing on engagement, such algorithms unwillingly intensify selective exposure and modeling behavioral proxies of user attention and selective exposure, but the process by which content curation is converted into different attention has not been empirically studied. In this research, a hybrid DistilBERT‐BiLSTM‐CatBoost model is suggested to quantitatively examine the joint effect of semantic features and behavioral history on user attention and cognition in an algorithmic feed. The framework works at the slate level using the Microsoft news dataset (MIND) that includes impressions level exposure logs and rich textual metadata, but relies on three complementary processes: (1) DistilBERT encodes the semantic content of news titles and abstracts; (2) BiLSTM takes into consideration the sequential user preferences and short term behavioral dynamics; (3) CatBoost is a bias aware learning with the capability to rank semantic content by means of textual, sequential, and contextual covariates, such as position, popularity, rec Counterfactual evaluation techniques (propensity weighted) (IPS weighted NDCG) are used to identify the real relevance and exposure effects. The experimental findings indicate that the proposed model is more superior as compared to five states of the art baselines (XGBoost, LightGBM, Random Forest, Decision Tree, and AdaBoost) with an F1 score of 0.956 and NDCG at 5 of 0.872. The effectiveness of the DistilBERT embedding in generalization is established by comparing across five deep architectures (LSTM, BiLSTM, GRU, CNN BiLSTM, and Transformer Encoder). SHAP attributed feature identification shows position and popularity to be the strongest predictors of attention, then semantic similarity and recency, the latter continuing to have a strong effect on user cognition despite bias correction. The results overall give a replicable, empirically‐driven model of bias conscious ranking analysis in personalized feeds, which provides insights on content algorithms in a scalable way in both maintaining engagement and contributing to the passive process of rationalizing modeling behavioral proxies of user attention and selective exposure.
Authors
- Chao Hua Yan (ORCID: https://orcid.org/0009-0004-9681-3539)
Institutions
- Nanyang Institute of Technology (CN)
Publication Details
- Journal
- Computational Intelligence
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1111/coin.70300
- Primary Topic
- Sentiment Analysis and Opinion Mining
- Type
- article
- Field-Weighted Citation Impact
- 0.00