Software Engineer Graduate (MLOps) - 2027 Start

TikTok Inc

  • San Jose, CA
  • 1 day ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Automationunmatched
    • Business Modelunmatched
    • C Programming Languageunmatched
    • C++ Programming Languageunmatched
    • CPU (Central Processing Unit)unmatched
    • Capacity Managementunmatched
    • Communication Skillsunmatched
    • Computer Scienceunmatched
    • Continuous Improvementunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Go Programming Language (Golang)unmatched
    • Identify Issuesunmatched
    • Javaunmatched
    • Linux Operating Systemunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Onboardingunmatched
    • Performance Tuning/Optimizationunmatched
    • Programming Languagesunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Managementunmatched
    • Resource Managementunmatched
    • Software Engineeringunmatched

    Description

    The monetization technology team works on building and running large-scale, globally distributed, fault-tolerant ads systems. SREs keep the systems up and running with the highest level of availability, ensuring our users have the best experience possible.

    We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

    Responsibilities:

    • Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services.
    • Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates
    • Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization.
    • Lead the enablement of new machine learning frameworks and GPU platforms, driving large-scale production deployment while maintaining model quality and business performance.
    • Design, build, and maintain MLOps platforms and automation tools, including quota management, job diagnostics, monitoring and observability, resource management, and operational tooling. Minimum Qualifications:
    • Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Computer Science or a related discipline.
    • Expertise in Linux operating systems, networking, storage.
    • Experience programming in at least one of the following programming languages: Python, Go, C, C++, or Java.
    • Experience in troubleshooting application issues, or production operations.
    • Effective communication skills and a sense of ownership and drive.

    Preferred Qualifications:

    • Experience in SRE of machine learning systems.
    • Experience in SRE of ads/recommendation/search systems.

    By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy

    Numbers & Facts

    LocationSan Jose, CA

    Similar Jobs