Tech Lead Cloud Site Reliability Engineer - DCS Cloud

Beijing ByteDance Technology Co Ltd

  • San Jose, CA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Automationunmatched
    • Blogunmatched
    • C++ Programming Languageunmatched
    • CUDA (Compute Unified Device Architecture)unmatched
    • Capacity Managementunmatched
    • Change Managementunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Networksunmatched
    • Computer Scienceunmatched
    • Corporate Complianceunmatched
    • Cost Controlunmatched
    • Customer Support/Serviceunmatched
    • DevOpsunmatched
    • Distributed Control Systems (DCS)unmatched
    • Dockerunmatched
    • Ecosystemsunmatched
    • GCP (Good Clinical Practices)unmatched
    • GPU (Graphics Processing Unit)unmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • Information/Data Security (InfoSec)unmatched
    • K Virtual Machine (KVM)unmatched
    • Large-Scale Systemsunmatched
    • Linux Operating Systemunmatched
    • Machine Toolunmatched
    • Microsoft Windows Azureunmatched
    • Network Operations Centerunmatched
    • Network Performance/Analysisunmatched
    • On Callunmatched
    • Open Sourceunmatched
    • Operating Systemsunmatched
    • Patentsunmatched
    • Private Cloudunmatched
    • Process Improvementunmatched
    • Product Lifecycleunmatched
    • Production Systemsunmatched
    • Programming Languagesunmatched
    • Public Cloudunmatched
    • Python Programming/Scripting Languageunmatched
    • Regulatory Complianceunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Software Engineeringunmatched
    • Stress Testingunmatched
    • System Operationsunmatched
    • Team Playerunmatched
    • Technical Leadershipunmatched
    • Technical Operationsunmatched
    • Testingunmatched
    • Topologyunmatched
    • Vehicle Fleetsunmatched
    • Virtualizationunmatched

    Description

    Our Infrastructure Engineering team supports the companys fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable.

    Responsibilities - What Youll Do

    • Design, build, scale, and operate ByteDance's global infrastructure, including large-scale systems spanning public and private clouds.
    • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
    • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the companys global compliance standards.
    • Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
    • Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

    Minimum Qualifications

    • Bachelor's degree or above in Computer Science, Software Engineering, Information Security, or a related field.
    • 5+ years of experience in Linux operations, SRE, or DevOps; experience operating large-scale production environments is a strong plus.
    • Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
    • Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills.
    • Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes.
    • Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset.

    Preferred Qualifications

    • Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms.
    • Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU.
    • Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces.
    • Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines.
    • Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization.
    • Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred.

    Numbers & Facts

    LocationSan Jose, CA

    Similar Jobs