Real-Time Pedestrian Crossing Intent Prediction and Risk Assessment Framework Using Skeleton Graph Convolutional Networks

Pedestrian safety at urban intersections remains a major challenge in Intelligent Transportation Systems (ITSs). This study investigates whether crossing intention can be reliably inferred directly from temporal body-pose dynamics to drive real-time collision warnings on embedded edge platforms. Existing vision-based approaches that rely primarily on bounding-box proximity or scene-level spatial grids are often prone to false alarms in complex urban environments with motorcycles, stationary pedestrians, and background clutter. To overcome these limitations, we propose an end-to-end framework consisting of four sequential processing stages: (1) a perception layer integrating YOLOv8s, ByteTrack, a displacement filter, and rider suppression to generate reliable pedestrian trajectories; (2) a skeleton extraction layer utilizing YOLOv8s-pose to construct temporal sequences of 17 anatomical keypoints; (3) an ultra-lightweight Skeleton Graph Convolutional Network (SkeletonGCN, comprising 33.8 K parameters, <0.2 MB) that models body-joint kinematics and temporal motion dynamics; and (4) an image-space Time-to-Collision (TTC) risk-fusion module. While this fusion approach avoids explicit geometric camera calibration, it still relies on predefined scene-profile parameters and image-space motion assumptions. Furthermore, while the intention classifier is quantitatively evaluated, the risk-fusion module is procedurally defined, and its resulting four-level collision warnings are demonstrated operationally rather than validated against ground-truth hazard annotations. Evaluated on 49,948 valid sequences from the JAAD and PIE benchmark datasets under a strict video-level partitioning protocol, the unified SkeletonGCN achieves a macro-F1 score of 0.717 (with per-scene subset macro-F1 scores of 0.761 on JAAD/PIE urban and 0.895 on intersections), significantly outperforming baseline models. When deployed on an NVIDIA Jetson Orin NX edge device using TensorRT FP16, the full pipeline achieves an instrumented latency of 70.7 ms per frame (~14 fps) and a sustained wall-clock throughput of 7.4 fps on real-world urban dashcam video. System limitations include sensitivity to 2D printed human imagery and reduced prediction reliability under low-light nighttime conditions.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-10
DOI
https://doi.org/10.3390/electronics15184106
Primary Topic
Autonomous Vehicle Technology and Safety
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Real-Time Pedestrian Crossing Intent Prediction and Risk Assessment Framework Using Skeleton Graph Convolutional Networks

Chayanon Sub-r-pa, Rung-Ching Chen, Yi-Xuan Deng
Electronics
Autonomous Vehicle Technology and Safety
article

Real-Time Pedestrian Crossing Intent Prediction and Risk Assessment Framework Using Skeleton Graph Convolutional Networks

Chayanon Sub-r-pa, Rung-Ching Chen, Yi-Xuan Deng
article en

Abstract

Pedestrian safety at urban intersections remains a major challenge in Intelligent Transportation Systems (ITSs). This study investigates whether crossing intention can be reliably inferred directly from temporal body-pose dynamics to drive real-time collision warnings on embedded edge platforms. Existing vision-based approaches that rely primarily on bounding-box proximity or scene-level spatial grids are often prone to false alarms in complex urban environments with motorcycles, stationary pedestrians, and background clutter. To overcome these limitations, we propose an end-to-end framework consisting of four sequential processing stages: (1) a perception layer integrating YOLOv8s, ByteTrack, a displacement filter, and rider suppression to generate reliable pedestrian trajectories; (2) a skeleton extraction layer utilizing YOLOv8s-pose to construct temporal sequences of 17 anatomical keypoints; (3) an ultra-lightweight Skeleton Graph Convolutional Network (SkeletonGCN, comprising 33.8 K parameters, <0.2 MB) that models body-joint kinematics and temporal motion dynamics; and (4) an image-space Time-to-Collision (TTC) risk-fusion module. While this fusion approach avoids explicit geometric camera calibration, it still relies on predefined scene-profile parameters and image-space motion assumptions. Furthermore, while the intention classifier is quantitatively evaluated, the risk-fusion module is procedurally defined, and its resulting four-level collision warnings are demonstrated operationally rather than validated against ground-truth hazard annotations. Evaluated on 49,948 valid sequences from the JAAD and PIE benchmark datasets under a strict video-level partitioning protocol, the unified SkeletonGCN achieves a macro-F1 score of 0.717 (with per-scene subset macro-F1 scores of 0.761 on JAAD/PIE urban and 0.895 on intersections), significantly outperforming baseline models. When deployed on an NVIDIA Jetson Orin NX edge device using TensorRT FP16, the full pipeline achieves an instrumented latency of 70.7 ms per frame (~14 fps) and a sustained wall-clock throughput of 7.4 fps on real-world urban dashcam video. System limitations include sensitivity to 2D printed human imagery and reduced prediction reliability under low-light nighttime conditions.

ElectronicsVol. 15(18)
Chaoyang University of Technology (TW)
Sustainable cities and communities
Openalex Percentile: Top 18%
Autonomous Vehicle Technology and Safety
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.