A REVIEW OF DEEP LEARNING AND SYNTHETIC DATA GENERATION FOR HUMAN POSE ESTIMATION IN AUTONOMOUS ROBOTIC SYSTEMS: METHODS, CHALLENGES, AND FUTURE DIRECTIONS
Keywords:
Human Pose Estimation, Deep Learning, Autonomous Robotic Systems, Synthetic Data, Synthetic-to-Real TransferAbstract
Reliable perception of human posture and movement is essential for autonomous robotic systems operating in shared or dynamic environments. Human pose estimation provides this capability by detecting body keypoints and reconstructing human motion from visual data, but its practical use in robotics is still affected by challenges such as occlusion, viewpoint changes, limited real-time performance, and insufficient annotated training data. This paper analyzes recent developments in 2D and 3D human pose estimation from the perspective of autonomous robotic perception. The review covers fundamental body representation models, single-person and multi-person estimation strategies, and the main differences between skeleton-based and model-based 3D approaches. Special attention is given to deep learning architectures, including convolutional neural networks, vision transformers, and emerging state-space models, with emphasis on their accuracy, computational complexity, and suitability for real-time robotic applications. In addition, the paper discusses synthetic data generation as a practical solution for improving dataset diversity and annotation quality. Rendering-based simulation, motion-capture data, generative methods, and fusion-based techniques are discussed as tools for producing controllable training samples. A small proof-of-concept example based on a game-based 3D environment is also included to demonstrate how camera and character positions can be extracted and used to automatically calculate camera-to-human distance as a spatial annotation. The main limitation of this approach is the synthetic-to-real domain gap, which can reduce model reliability when synthetic datasets are applied to real-world robotic scenarios. Future research should therefore focus on robot-oriented synthetic datasets, improved synthetic-to-real transfer, generative artificial intelligence, and multimodal sensing to support more robust human perception in autonomous robotic systems.