The mission of our AML team is to push next-generation recommendation-based algorithms and platform for the company. We also drive substantial impact for core businesses of the company. Currently we are looking for Production Engineers to join our team to support and advance that mission
Responsibilities:
System Stability & Production Management: Responsible for the production management and stability assurance of AML (Applied Machine Learning) training, inference, and storage systems. This covers core pipelines including scheduling and orchestration, K8s/GPU clusters, distributed training, online inference serving, and ParameterServer/NoSQL storage.
Reliability Engineering: Build and maintain mechanisms for SLO/SLA, observability, alerting, On-call processes, fault diagnosis, auto-healing, disaster recovery, and incident reviews (post-mortems).
Engineering Excellence: Drive engineering capabilities such as CI/CD, canary releases, auto-rollback, automated inspections, pre-flight checks, capacity forecasting, and elastic auto-scaling.
Resource & Cost Management: Oversee resource governance across GPU/CPU/storage/network, including quota management, cost attribution, and performance tuning, to improve system availability, resource utilization, and overall R&D efficiency.Minimum Qualification(s):
Bachelor's degree or above in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
Familiar with Linux and proficient in at least one of the following programming/scripting languages: Shell, Python, Go, or C++.
Understanding of machine learning training/inference architectures, Kubernetes, GPU clusters, or distributed storage systems.
Proven experience in online troubleshooting, performance analysis, and building automation platforms.
Strong sense of responsibility, clear logical thinking, and the ability to drive the resolution of complex issues across cross-functional teams.
Preferred Qualification(s):
Experience with large-scale training/inference/storage platforms, SLO governance, FinOps, NoSQL, or open-source infrastructure is highly preferred.
Numbers & Facts
Location
San Jose, CA
Skills
Algorithmsunmatched
Artificial Intelligence (AI)unmatched
Autoscalingunmatched
C++ Programming Languageunmatched
CPU (Central Processing Unit)unmatched
Computer Scienceunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Cost Controlunmatched
Cross-Functionalunmatched
Disaster Recoveryunmatched
Distributed Computingunmatched
Fault Managementunmatched
Forecastingunmatched
GPU (Graphics Processing Unit)unmatched
Go Programming Language (Golang)unmatched
Home Automationunmatched
Identify Issuesunmatched
Incident Managementunmatched
Linux Operating Systemunmatched
Machine Learningunmatched
Mail Processingunmatched
NoSQLunmatched
On Callunmatched
Online Trainingunmatched
Open Sourceunmatched
Performance Analysisunmatched
Performance Tuning/Optimizationunmatched
Problem Solving Skillsunmatched
Production Managementunmatched
Python Programming/Scripting Languageunmatched
Reliability Engineeringunmatched
Research & Development (R&D)unmatched
Resource Managementunmatched
Resource Utilizationunmatched
Scripting (Scripting Languages)unmatched
Service Level Agreement (SLA)unmatched
Software Engineeringunmatched
Unix Shell Programmingunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.