Amazon.com Inc logo

Cloud Hardware Development Manager, AWS Gen AI & ML Servers

Amazon.com Inc
  • Cupertino, CA
    7 days ago

    Job Description

    AWS operates the world"s largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms - from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

    We are seeking a Cloud Hardware Development Manager to lead a team of hardware, systems development engineers, and technical program managers responsible for the design, validation, manufacturing, and fleet operations of GPU-accelerated server platforms. You will set technical direction for your team, manage ODM and silicon supplier partnerships, drive hardware programs from concept through production ramp, and own fleet reliability metrics. This role requires deep hardware expertise to guide technical decisions, combined with people leadership to hire, develop, and retain a high-performing engineering team.

    What You Will Do

    You will lead hardware engineers who define the platforms running the world"s largest AI workloads. You will set technical strategy for your team"s hardware portfolio working alongside product and leadership teams, making trade-offs between schedule, cost, reliability, and performance. You will guide your engineers through ambiguous design problems while removing blockers and ensuring delivery. You will own the end-to-end lifecycle of your team"s hardware, from architecture definition through fleet operations, and be accountable for fleet quality metrics including annualized failure rates and system availability.

    Why You Will Love It

    The world"s most advanced frontier models train on the hardware your team designs. You will build and lead the engineers who define GPU server platforms at unprecedented scale. Your decisions on architecture, components, and quality directly impact fleet reliability for customers within weeks of deployment. The team is deeply technical and high-trust, with direct access to senior leadership and the freedom to set technical direction.

    The Ideal Candidate

    You combine deep hardware expertise with strong people leadership. You have designed or led teams that delivered server hardware at scale, and you can still dive into a schematic review or signal integrity issue when needed. You set high standards for your team and your partners, make data-driven decisions, and communicate clearly to both engineers and executives. You build teams that are stronger because of your presence but don"t require it to be successful.

    Key job responsibilities

    Technical Leadership & Strategy

    • Set technical direction for your team"s hardware portfolio across GPU-accelerated server platforms, making architecture and component trade-offs aligned with customer requirements and business goals
    • Guide system-level design decisions across thermal, mechanical, power delivery, signal integrity, and accelerator subsystems - including trade-offs on cooling architecture (liquid vs. air), power budget allocation, and PCIe/interconnect topology
    • Drive design reviews, qualification gates, and go/no-go decisions with deep technical judgment
    • Own fleet reliability and availability for your team"s hardware; drive continuous improvement to reduce failure rates

    People Leadership & Development

    • Hire, develop, and retain a team of hardware and system software engineers; build a high-performing team culture focused on engineering excellence and customer obsession
    • Set clear goals, provide regular feedback, and actively coach engineers through career growth as part of 1:1s
    • Perform calibrations and promotion assessments; ensure your team"s technical bar remains high
    • Create an inclusive team environment where engineers can do their best work

    Program Delivery & ODM Management

    • Drive hardware programs from concept through manufacturing and fleet deployment, ensuring milestones are met across NPI phases (Program Initiation, Design, Qualification, Pilot, Post-GA)
    • Lead ODM and silicon supplier partnerships: establish technical standards, drive design reviews, manage EVT/DVT/PVT builds, and hold partners accountable for quality and schedule
    • Identify program risks early, escalate with data and proposed mitigations, and unblock cross-functional dependencies

    Fleet Operations & Quality

    • Own operational metrics for your team"s server platforms: annualized failure rates, system availability, manufacturing escape rates
    • Drive root cause analysis of fleet-wide failures and ensure corrective actions flow back into design requirements and qualification criteria for future platforms
    • Establish closed-loop feedback systems connecting field failure data to upstream design and manufacturing improvements

    Cross-Functional Alignment

    • Align with EC2 architecture teams on instance requirements, workload characterization, and platform roadmaps
    • Partner with firmware, software, test automation, and datacenter operations teams to ensure hardware is debuggable, serviceable, and automation-ready
    • Communicate technical strategy, program status, and risk posture to senior leadership

    May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.

    A day in the life

    You start the day in a 1:1 with one of your engineers, coaching them through a design trade-off on power delivery for a next-gen accelerator platform. Mid-morning, you lead a design review with your ODM partner, driving closure on open signal integrity items from the DVT build. In the afternoon, you review fleet reliability data with your team, identifying a component failure trend and aligning on corrective actions. You end the day in a planning session with senior leadership, presenting your team"s hardware roadmap and resource needs for the next platform generation.

    About the team

    The Hardware Engineering AI/ML Ultraserver platform team is a group of engineers and technical program managers directly responsible for launching and maintaining GPU-accelerated servers in the AWS fleet. Located in Seattle, Cupertino, and Austin, we work with internal engineering teams, ODMs, and design partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams.

    Numbers & Facts

    LocationCupertino, CA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Skills

    • Alliance/Partner Managementunmatched
    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Amazon Web Services (AWS)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Budgetingunmatched
    • Career Counselingunmatched
    • Cloud Computingunmatched
    • Coachingunmatched
    • Communication Skillsunmatched
    • Component Selectionunmatched
    • Computer Engineeringunmatched
    • Computer Firmwareunmatched
    • Continuous Improvementunmatched
    • Corrective Actionunmatched
    • Cross-Functionalunmatched
    • Field Trialsunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Hardware Architectureunmatched
    • Hardware Designunmatched
    • Hardware Developmentunmatched
    • Leadershipunmatched
    • Manufacturingunmatched
    • Manufacturing Operationsunmatched
    • Manufacturing Systemsunmatched
    • Metricsunmatched
    • Network Operations Centerunmatched
    • Original Design Manufacturer (ODM)unmatched
    • PCI Express (PCI-E)unmatched
    • Process Improvementunmatched
    • Product/Service Launchunmatched
    • Project/Program Managementunmatched
    • Quality Metricsunmatched
    • Requirements Managementunmatched
    • Riskunmatched
    • Risk Analysisunmatched
    • Root Cause Analysisunmatched
    • Schematicsunmatched
    • Server Architectureunmatched
    • Server Hardwareunmatched
    • Set Goalsunmatched
    • Signal Integrityunmatched
    • Software Engineeringunmatched
    • Software Testingunmatched
    • Strategic Planningunmatched
    • Systems Engineeringunmatched
    • Team Lead/Managerunmatched
    • Team Playerunmatched
    • Technical Leadershipunmatched
    • Technical Strategyunmatched
    • Test Automationunmatched
    • Topologyunmatched
    • Vehicle Fleetsunmatched
    • Willing to Travelunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder