Want to know if you’re a fit? Upload your resume and let our AI show you.
Skills
Automationunmatched
Business Modelunmatched
C Programming Languageunmatched
C++ Programming Languageunmatched
CPU (Central Processing Unit)unmatched
Capacity Managementunmatched
Communication Skillsunmatched
Computer Scienceunmatched
Continuous Improvementunmatched
GPU (Graphics Processing Unit)unmatched
Go Programming Language (Golang)unmatched
Identify Issuesunmatched
Javaunmatched
Linux Operating Systemunmatched
Machine Learningunmatched
Machine Toolunmatched
Onboardingunmatched
Performance Tuning/Optimizationunmatched
Programming Languagesunmatched
Python Programming/Scripting Languageunmatched
Quality Managementunmatched
Resource Managementunmatched
Software Engineeringunmatched
Description
The monetization technology team works on building and running large-scale, globally distributed, fault-tolerant ads systems. SREs keep the systems up and running with the highest level of availability, ensuring our users have the best experience possible.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities:
Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services.
Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates
Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization.
Lead the enablement of new machine learning frameworks and GPU platforms, driving large-scale production deployment while maintaining model quality and business performance.
Design, build, and maintain MLOps platforms and automation tools, including quota management, job diagnostics, monitoring and observability, resource management, and operational tooling. Minimum Qualifications:
Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Computer Science or a related discipline.
Expertise in Linux operating systems, networking, storage.
Experience programming in at least one of the following programming languages: Python, Go, C, C++, or Java.
Experience in troubleshooting application issues, or production operations.
Effective communication skills and a sense of ownership and drive.
Preferred Qualifications:
Experience in SRE of machine learning systems.
Experience in SRE of ads/recommendation/search systems.
By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy