Real-Time Pedestrian Crossing Intent Prediction and Risk Assessment Framework Using Skeleton Graph Convolutional Networks
Pedestrian safety at urban intersections remains a major challenge in Intelligent Transportation Systems (ITSs). This study investigates whether crossing intention can be reliably inferred directly from temporal body-pose dynamics to drive real-time collision warnings on embedded edge platforms. Existing vision-based approaches that rely primarily on bounding-box proximity or scene-level spatial grids are often prone to false alarms in complex urban environments with motorcycles, stationary pedestrians, and background clutter. To overcome these limitations, we propose an end-to-end framework consisting of four sequential processing stages: (1) a perception layer integrating YOLOv8s, ByteTrack, a displacement filter, and rider suppression to generate reliable pedestrian trajectories; (2) a skeleton extraction layer utilizing YOLOv8s-pose to construct temporal sequences of 17 anatomical keypoints; (3) an ultra-lightweight Skeleton Graph Convolutional Network (SkeletonGCN, comprising 33.8 K parameters, <0.2 MB) that models body-joint kinematics and temporal motion dynamics; and (4) an image-space Time-to-Collision (TTC) risk-fusion module. While this fusion approach avoids explicit geometric camera calibration, it still relies on predefined scene-profile parameters and image-space motion assumptions. Furthermore, while the intention classifier is quantitatively evaluated, the risk-fusion module is procedurally defined, and its resulting four-level collision warnings are demonstrated operationally rather than validated against ground-truth hazard annotations. Evaluated on 49,948 valid sequences from the JAAD and PIE benchmark datasets under a strict video-level partitioning protocol, the unified SkeletonGCN achieves a macro-F1 score of 0.717 (with per-scene subset macro-F1 scores of 0.761 on JAAD/PIE urban and 0.895 on intersections), significantly outperforming baseline models. When deployed on an NVIDIA Jetson Orin NX edge device using TensorRT FP16, the full pipeline achieves an instrumented latency of 70.7 ms per frame (~14 fps) and a sustained wall-clock throughput of 7.4 fps on real-world urban dashcam video. System limitations include sensitivity to 2D printed human imagery and reduced prediction reliability under low-light nighttime conditions.
Authors
- Chayanon Sub-r-pa (ORCID: https://orcid.org/0000-0001-7345-9574)
- Rung-Ching Chen (ORCID: https://orcid.org/0000-0001-7621-1988)
- Yi-Xuan Deng
Institutions
- Chaoyang University of Technology (TW)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/electronics15184106
- Primary Topic
- Autonomous Vehicle Technology and Safety
- Type
- article
- Field-Weighted Citation Impact
- 0.00