Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China's large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services.
Data AML is ByteDance's machine learning platform team. It provides training and inference systems for recommendation, advertising, computer vision, speech, and NLP scenarios across products such as Douyin, Toutiao, and Xigua Video. The team also supports internal business teams with large-scale machine learning compute, explores general and innovative algorithms for business problems, and offers core machine learning and recommendation system capabilities to external enterprise customers through Volcano Engine.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities:
- Design and develop resource scheduling systems for machine learning workloads, supporting Volcano Ark and machine learning platform products.
- Optimize orchestration and scheduling of heterogeneous compute resources, including GPUs, CPUs, and other accelerators, as well as storage resources such as cloud storage and networking resources such as VPC and RDMA, across multiple data centers and clusters.
- Support scheduling requirements for offline training, online inference, and other workloads under strict multi-tenant isolation, improving overall resource utilization and efficiency.Minimum Qualifications:
- Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science or a related discipline.
- Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
- Solid foundation in computer science and programming, familiarity with common algorithms and data structures, and good coding habits.
- Familiar with at least one mainstream machine learning framework, such as TensorFlow, PyTorch, or an internally developed framework.
- Familiar with Kubernetes architecture and ecosystem, as well as container technologies such as Docker, container, and Kata;
- Understands distributed system principles and has participated in the design, development, or maintenance of large-scale distributed systems.
Preferred Qualifications:
- Practical experience in large-scale cluster online/offline resource scheduling; source-level understanding of one or more open-source schedulers such as Kubernetes, Volcano, YARN, or Mesos; familiarity with containerization and lightweight virtualization technologies.
- Deep understanding and practical experience in scheduling topics such as multi-tenant quota governance, preemption, elasticity, fragmentation, tidal scheduling, co-location, and QoS; strong analytical and modeling ability for complex problems; GPU scheduling experience is preferred.
- Experience in at least one of the following areas: CUDA, RDMA, AI infrastructure, hardware and software co-design, high-performance computing, machine learning hardware architecture such as GPUs, accelerators and networking, ML for systems, or distributed storage.
- Hands-on experience in cloud-native machine learning systems is preferred.