Site Reliability Engineer - System Service Global

Beijing ByteDance Technology Co Ltd
  • San Jose, CA
    30+ days ago

    Job Description

    The Global System Service team owns the infrastructure services and management solutions that power ByteDance's data centers outside of China from day-to-day operations to long-term architecture design and maintenance. The team specializes in composing end-to-end solutions by drawing on both open-source community tools and in-house developed products, tailored to both the business requirements and the operational complexities of large-scale infrastructure across ByteDance's non-China regions. Our mission is to deliver efficient infrastructure solutions and a stable, secure system environment for ByteDance's global business.

    We are looking for a self-motivated system engineer that is equipped with SRE mindset and DevOps skills. Your responsibilities will include:

    • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring.
    • Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
    • Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions.
    • Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortem reviews to drive continuous reliability improvements.
    • Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth.
    • Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.Minimum Qualifications:
    • Bachelors degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
    • Solid experience in large-scale Linux host management, including OS deployment, configuration management, patching, and fleet operations.
    • Strong hands-on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
    • Proficiency with DevOps tooling, including configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines.
    • Familiarity with SRE principles and practices, including SLO/SLI definition, error budget management, and blameless post-mortems.
    • Solid understanding of high availability design patterns, active-active/active-passive architectures, and disaster recovery strategies.
    • Strong troubleshooting skills across the Linux system stack and network layer.

    Preferred Qualifications:

    • Experience managing host fleets at scale (thousands of nodes or above) in a production environment.
    • Scripting or development experience in Python, Go, or Bash for automation and tooling.
    • Exposure to hybrid or multi-region data center environments.

    Numbers & Facts

    LocationSan Jose, CA

    Skills

    • Ansibleunmatched
    • Applications Securityunmatched
    • Architectural Servicesunmatched
    • Authenticationunmatched
    • Automationunmatched
    • BIND (Berkeley Internet Name Domain) DNS Server softwareunmatched
    • Bash Scriptingunmatched
    • Budget Managementunmatched
    • Business Growthunmatched
    • Computer Engineeringunmatched
    • Computer Scienceunmatched
    • Configuration Managementunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • DHCP (Dynamic Host Configuration Protocol)unmatched
    • DNS (Domain Name System)unmatched
    • Design Patterns Programming Methodologiesunmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Electrical Engineeringunmatched
    • Fleet Managementunmatched
    • Go Programming Language (Golang)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • International Businessunmatched
    • Kerberosunmatched
    • Linux Operating Systemunmatched
    • Machine Toolunmatched
    • NAT (Network Address Translation)unmatched
    • NTPunmatched
    • Network Operations Centerunmatched
    • Network Securityunmatched
    • Open Sourceunmatched
    • Operating Systemsunmatched
    • Operational Strategyunmatched
    • Process Improvementunmatched
    • Product Developmentunmatched
    • Production Systemsunmatched
    • Puppet (Configuration Management)unmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Patchesunmatched
    • Systems Engineeringunmatched
    • Vehicle Fleetsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder