We are looking for a Machine Learning Engineer to join our core research and development team focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.
Human demonstration data is the foundation of robot learning, and its quality depends on accurately reconstructing human motion. In this role, you will develop models and production pipelines that transform head-mounted and body-mounted camera streams—including wide-FOV, stereo, motion-blurred, and heavily self-occluded video—into metrically accurate, temporally consistent 3D pose representations for robot policy training and human-to-robot motion retargeting.
You will work across the entire perception stack, including camera calibration, data annotation, model training, evaluation, and large-scale deployment. This role is ideal for engineers with strong expertise in both computer vision and deep learning who enjoy solving challenging real-world perception problems.
Responsibilities
Develop state-of-the-art 3D body and hand pose estimation models for egocentric video using monocular and stereo camera systems.
Build models for 2D/3D keypoint estimation, SMPL/SMPL-X, MANO, and full-body motion reconstruction.
Address challenging egocentric vision problems, including severe self-occlusion, motion blur, rolling shutter artifacts, truncated limbs, extreme viewpoints, and hand-object interaction.
Design and maintain camera geometry and calibration pipelines, including fisheye and wide-FOV camera models, stereo calibration, triangulation, and coordinate frame alignment.
Improve temporal consistency and physical plausibility using filtering, kinematic constraints, multi-view fusion, and multi-modal sensor integration.
Build scalable annotation, evaluation, and quality assurance pipelines for large-scale human motion datasets.
Optimize large-scale model training and high-throughput inference pipelines for production environments.
Collaborate closely with robotics engineers to convert reconstructed human motion into high-quality robot training data.
Contribute to system architecture, engineering best practices, and the long-term evolution of the perception platform.
Minimum Qualifications
Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related field.
3+ years of experience building and deploying machine learning systems.
Hands-on experience with 3D human pose estimation, hand pose estimation, or human motion tracking from video.
Strong understanding of multi-view geometry, camera calibration, triangulation, coordinate transformations, and projection models.
Strong Python programming skills and proficiency with PyTorch or TensorFlow.
Solid knowledge of modern deep learning techniques, model training, evaluation, and production ML workflows.
Strong analytical and problem-solving skills with the ability to thrive in a fast-paced collaborative environment.
Preferred Qualifications
Experience with egocentric perception systems, AR/VR headsets, smart glasses, or wearable capture rigs.
Expertise with SMPL, SMPL-X, MANO, inverse kinematics, markerless motion capture, or hand-object pose estimation.
Familiarity with egocentric vision datasets and benchmarks.
Experience with modern video and 3D learning architectures, including Video Transformers, diffusion-based motion models, or 3D CNNs.
Experience building multi-camera capture systems, synchronization, and calibration infrastructure.
Experience with human-to-robot motion retargeting, teleoperation, imitation learning, or dexterous manipulation.
Experience developing annotation tools, active learning pipelines, or large-scale data quality systems.
Publications at leading conferences such as CVPR, ICCV, ECCV, NeurIPS, SIGGRAPH, or 3DV, open-source contributions, or demonstrated impact in applied AI systems.
What We Offer
Competitive salary and equity package
Comprehensive medical, dental, and vision insurance
401(k) retirement plan
Generous paid time off and company holidays
Paid sick leave
Opportunity to work alongside leading researchers and engineers in robotics, computer vision, and AI
High-impact role building cutting-edge perception systems for next-generation robotics
Numbers & Facts
Location
Santa Clara, California
Website
https://www.glinttechsolutions.com
Skills
3D Modelingunmatched
Accounts Receivableunmatched
Analysis Skillsunmatched
Artificial Intelligence (AI)unmatched
Benchmarkingunmatched
Best Practicesunmatched
Calibrationunmatched
Computer Programmingunmatched
Computer Scienceunmatched
Computer Visionunmatched
Conferencesunmatched
Data Modelingunmatched
Data Qualityunmatched
Data Setsunmatched
Deep Learningunmatched
FishEyeunmatched
Geometryunmatched
High Throughputunmatched
Large-Scale Systemsunmatched
Machine Learningunmatched
Metricsunmatched
Open Sourceunmatched
Problem Solving Skillsunmatched
Production Systemsunmatched
Publicationsunmatched
Python Programming/Scripting Languageunmatched
Quality Assuranceunmatched
Research & Development (R&D)unmatched
Roboticsunmatched
Scalable System Developmentunmatched
System Architectureunmatched
Team Playerunmatched
Training Data Setsunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.