We are hiring a Software Engineer to productionize and optimize our GPU serving stack, working across our custom inference APIs, the vLLM serving runtime, the AMD ROCm software stack, and rack-scale AMD GPU infrastructure, to make this new serving path reliable, numerically correct, observable, and exceptionally performant. Design, build, deploy, and maintain the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.