Temporal Consistency-Constrained Vision–Action Latent Space Alignment and Policy Optimization

Real-world visuomotor control requires policies to reason over uncertain visual observations, action continuity, and long-horizon error accumulation. Existing vision-action policies often improve action generation but insufficiently align visual state evolution with action-induced latent dynamics. This paper proposes TCVA-PO, a temporal consistencyconstrained framework for vision-action latent space alignment and policy optimization. Its defining contribution is the two-stage coupling of cross-modal latent increments from paired demonstrations with a policy-induced next-latent transition penalty; pointwise alignment, reconstruction, critic regression, advantage weighting, and smoothness supply complementary representation and optimization structure. Experiments on six public robot manipulation datasets show that TCVA-PO consistently improves task performance, latent consistency, action smoothness, and long-horizon robustness over representative imitation and generative visuomotor baselines under heterogeneous visual-control conditions and noisy demonstration settings. These results demonstrate that temporally aligned latent dynamics provide a more stable basis for robust policy learning in complex uncertain robotic systems.

Authors

Publication Details

Journal
International Journal of Pattern Recognition and Artificial Intelligence
Published
2026-09-21
DOI
https://doi.org/10.1142/s0218001426400641
Primary Topic
Robot Manipulation and Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Temporal Consistency-Constrained Vision–Action Latent Space Alignment and Policy Optimization

Yue Peng, Lei Zhang
International Journal of Pattern Recognition and Artificial Intelligence
Robot Manipulation and Learning
article

Temporal Consistency-Constrained Vision–Action Latent Space Alignment and Policy Optimization

Yue Peng, Lei Zhang
article en

Abstract

Real-world visuomotor control requires policies to reason over uncertain visual observations, action continuity, and long-horizon error accumulation. Existing vision-action policies often improve action generation but insufficiently align visual state evolution with action-induced latent dynamics. This paper proposes TCVA-PO, a temporal consistencyconstrained framework for vision-action latent space alignment and policy optimization. Its defining contribution is the two-stage coupling of cross-modal latent increments from paired demonstrations with a policy-induced next-latent transition penalty; pointwise alignment, reconstruction, critic regression, advantage weighting, and smoothness supply complementary representation and optimization structure. Experiments on six public robot manipulation datasets show that TCVA-PO consistently improves task performance, latent consistency, action smoothness, and long-horizon robustness over representative imitation and generative visuomotor baselines under heterogeneous visual-control conditions and noisy demonstration settings. These results demonstrate that temporally aligned latent dynamics provide a more stable basis for robust policy learning in complex uncertain robotic systems.

International Journal of Pattern Recognition and Artificial Intelligence
Peace, Justice and strong institutions
Openalex Percentile: Top 15%
Robot Manipulation and Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Temporal Consistency-Constrained Vision–Action Latent Space Alignment and Policy Optimization — Yue Peng, Lei Zhang · International Journal of Pattern Recognition and Artificial Intelligence (2026) | TGRS Research Map | TGRS