Trace Data of Secondary Students and Linguistic Analysis to Predict Learner Performance in a Multi‐Text Writing Task

ABSTRACT Background Writing from multiple sources is a cognitively demanding task that requires students to integrate reading, planning, and composition processes. While prior research has demonstrated the importance of self‐regulated learning (SRL) and linguistic features in predicting writing performance, most studies have focused on university students, with limited research targeting secondary school learners, particularly across multilingual contexts. Objective This study examined whether SRL processes derived from trace data and linguistic features from essays as learning products can be used to predict writing performance among secondary students from different countries and language backgrounds. Methods A total of 257 secondary students from Australia, Colombia, Germany, and Finland completed a multi‐source essay task in their school language within a digital learning environment. SRL processes were inferred from interaction trace data, while essay features—topic coverage, cohesion, and word count—were extracted using language‐specific analysis models. These data sources were then used to train several machine learning models to predict students' performance on essay writing. Results Among the machine learning models tested, the Random Forest classifier achieved the highest accuracy (0.80) and recall (0.71). The most influential predictors were elaboration/organisation (process), topic coverage and cohesion (product), and country. Results revealed distinct engagement patterns across countries, suggesting that language and educational context were associated with differences in both SRL processes and writing quality. Conclusions This study extends prior work by demonstrating that combining process and product features can support the prediction of writing performance across multilingual secondary education contexts. Findings emphasise the need for further research on younger learners and the role of linguistic background in SRL engagement.

Authors

Institutions

Publication Details

Journal
Journal of Computer Assisted Learning
Published
2026-09-21
DOI
https://doi.org/10.1002/jcal.70331
Primary Topic
Writing and Handwriting Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Trace Data of Secondary Students and Linguistic Analysis to Predict Learner Performance in a Multi‐Text Writing Task

Sadia Nawaz, Joni Lämsä, Dragan Gašević, Lyn Lim et al.
Journal of Computer Assisted Learning
Writing and Handwriting Education
article

Trace Data of Secondary Students and Linguistic Analysis to Predict Learner Performance in a Multi‐Text Writing Task

Sadia Nawaz, Joni Lämsä, Dragan Gašević, Lyn Lim, Maria Bannert, Sanna Järvelä, Mladen Raković, Gloria Milena Fernandez-Nieto, Xinyu Li, Keyang Qian, Halima Alnashiri
article en

Abstract

ABSTRACT Background Writing from multiple sources is a cognitively demanding task that requires students to integrate reading, planning, and composition processes. While prior research has demonstrated the importance of self‐regulated learning (SRL) and linguistic features in predicting writing performance, most studies have focused on university students, with limited research targeting secondary school learners, particularly across multilingual contexts. Objective This study examined whether SRL processes derived from trace data and linguistic features from essays as learning products can be used to predict writing performance among secondary students from different countries and language backgrounds. Methods A total of 257 secondary students from Australia, Colombia, Germany, and Finland completed a multi‐source essay task in their school language within a digital learning environment. SRL processes were inferred from interaction trace data, while essay features—topic coverage, cohesion, and word count—were extracted using language‐specific analysis models. These data sources were then used to train several machine learning models to predict students' performance on essay writing. Results Among the machine learning models tested, the Random Forest classifier achieved the highest accuracy (0.80) and recall (0.71). The most influential predictors were elaboration/organisation (process), topic coverage and cohesion (product), and country. Results revealed distinct engagement patterns across countries, suggesting that language and educational context were associated with differences in both SRL processes and writing quality. Conclusions This study extends prior work by demonstrating that combining process and product features can support the prediction of writing performance across multilingual secondary education contexts. Findings emphasise the need for further research on younger learners and the role of linguistic background in SRL engagement.

Journal of Computer Assisted LearningVol. 42(5)
Umm al-Qura University (SA), Education University of Hong Kong (HK), Monash University (AU), Technical University of Munich (DE), University of Hong Kong (HK), University of Oulu (FI)
Quality Education
Openalex Percentile: Top 3%
Writing and Handwriting Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.