Cloud Site Reliability Engineer - DCS Cloud

Beijing ByteDance Technology Co Ltd
  • Seattle, WA
    30+ days ago

    Job Description

    Our Infrastructure Engineering team supports the companys fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable.

    Responsibilities - What Youll Do

    • Design, build, scale, and operate ByteDance's global infrastructure, including large-scale systems spanning public and private clouds.
    • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
    • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the companys global compliance standards.
    • Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
    • Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

    Minimum Qualifications

    • Bachelor's degree or above in Computer Science, Software Engineering, Information Security, or a related field.
    • 2+ years of experience in Linux operations, SRE, or DevOps; experience operating large-scale production environments is a strong plus.
    • Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
    • Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills.
    • Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes.
    • Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset.

    Preferred Qualifications

    • Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms.
    • Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU.
    • Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces.
    • Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines.
    • Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization.
    • Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred.

    Numbers & Facts

    LocationSeattle, WA

    More jobs like this

    See more jobs

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.