Responsibilities: Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services; Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates; Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization; Lead the enablement of new machine learning frameworks and GPU platforms, driving large-scale production deployment while maintaining model quality and business performance; Design, build, and maintain MLOps platforms and automation tools, including quota management, job diagnostics, monitoring and observability, resource management, and operational tooling. Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.