About the Team
The ByteDance DPU (Data Processing Unit) team is building the foundational computing infrastructure for ByteDance and Volcano Engine Public Cloud. Our mission is to advance the architecture, development, and research of next-generation software-hardware technologies across compute, networking, and storage for cloud and AI computing. Our technology stack spans
GPU virtualization and scheduling for AI/ML workloads
We work at the intersection of software systems, distributed infrastructure, and custom hardware acceleration, shaping the next wave of cloud-scale computing.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
Design and develop DPU network software with a focus on high performance, low latency, and reliability.
Collaborate with hardware teams to build software-hardware co-design solutions for networking and storage acceleration.
Explore AI/ML infrastructure acceleration, leveraging DPUs, GPUs, and custom hardware to optimize distributed training and inference.
Drive end-to-end performance optimization, from OS kernels and drivers to user-space runtime systems.
Contribute to architecture design, technical proposals, and long-term research directions.Minimum Qualifications
Individuals who are completing or have recently completed a PhD degree in CS or a related discipline.
Proficiency in C/C++ development and debugging.
Familiar with Linux systems development experience.
Solid understanding of compute, network architecture, and operating systems.
Background in at least one of: software-hardware co-design, distributed systems, high-performance networking, or AI/ML systems.
Preferred Qualifications
Experience with software-hardware co-design (networking, storage, or distributed compute).
Hands-on experience with network virtualization (OVS, SR-IOV, eBPF).
Familiarity with DPDK and high-performance user-space networking.
Bonus points for hardware acceleration experience, FPGA/ASIC/GPU/CUDA
Bonus points for experience with NCCL Collectives along with AI communication patterns and parallelization techniques
Proven experience designing and building AI/ML infrastructure related but not limited to inference kv cache system, data preprocessing system.
Numbers & Facts
Location
San Jose, CA
Skills
ASIC (Application Specific Integrated Circuit)unmatched
Architectural Designunmatched
Artificial Intelligence (AI)unmatched
C Programming Languageunmatched
C++ Programming Languageunmatched
CUDA (Compute Unified Device Architecture)unmatched
Cloud Computingunmatched
Cloud Storageunmatched
Computer Architectureunmatched
Computer Hardwareunmatched
Computer Networksunmatched
Debugging Skillsunmatched
Distributed Computingunmatched
FPGAunmatched
GPU (Graphics Processing Unit)unmatched
Hardware Designunmatched
Hypervisorsunmatched
Kernel Programmingunmatched
Linux Operating Systemunmatched
Network Architecture/Engineeringunmatched
Network Designunmatched
Network Protocolsunmatched
Network Softwareunmatched
Onboardingunmatched
Operating Systemsunmatched
Performance Tuning/Optimizationunmatched
Scientific Researchunmatched
Software Designunmatched
Software Developmentunmatched
Technical/Engineering Designunmatched
Virtualizationunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.