In this role, you will develop models and production pipelines that transform head-mounted and body-mounted camera streams—including wide-FOV, stereo, motion-blurred, and heavily self-occluded video—into metrically accurate, temporally consistent 3D pose representations for robot policy training and human-to-robot motion retargeting. Address challenging egocentric vision problems, including severe self-occlusion, motion blur, rolling shutter artifacts, truncated limbs, extreme viewpoints, and hand-object interaction.