ByteBrain is ByteDance's AI for Infrastructure (AI4Infra) platform, dedicated to improving the efficiency, reliability, and intelligence of large-scale infrastructure systems through AI and machine learning. ByteBrain supports a wide range of infrastructure domains, including AI data center supply chains, AIOps, Operations Research and AgentOps, powering infrastructure optimization at massive scale.
Why This Role Is Unique
This role sits at the intersection of: Operations Research × AIOps × AI for Infra
You will have the opportunity to solve some of the most challenging optimization problems behind large-scale AI datacenters while pioneering the next generation of AI-powered decision-making systems, where LLMs, and optimization algorithms work together to improve efficiency, resource utilization, and operational intelligence across ByteDance's global infrastructure.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
- Design and develop AI, machine learning, and optimization algorithms to improve the efficiency, reliability, and performance of large-scale infrastructure systems and AI supply chain. Areas may include AIOps, operations research, software engineering, AgentOps, and system optimization.
- Drive the deployment, scaling, and continuous improvement of algorithms in production environments, supporting large-scale services.
- Identify optimization opportunities and emerging challenges from real-world infrastructure scenarios, translating them into impactful research and engineering solutions.
- Conduct cutting-edge research and publish high-quality papers in top-tier conferences and journals.Minimum Qualifications
- Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
- Proven research track record with multiple publications in top-tier conferences or journals related to AI, machine learning, operations research, systems, or related fields.
- Deep expertise in AI, machine learning, and/or operations research, with hands-on experience in large-scale data analysis and algorithm development.
- Strong coding, implementation, and problem-solving skills, with the ability to bridge research and production systems.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
- Industry experience applying AI and optimization techniques to real-world infrastructure challenges, such as:
- AI data center supply chain optimization
- AI Ops and intelligent operations
- Software engineering productivity optimization
- Operations research and resource scheduling
- System tuning and performance optimization
- Large-scale infrastructure management and automation