Amazon.com Inc logo

Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Amazon.com Inc
  • Cupertino, CA
    7 days ago

    Job Description

    Application deadline: Sep 1, 2026

    Do you want to build the infrastructure that keeps Artificial Intelligence compute capacity available to Generative AI customers? Do you want to solve problems at the boundary between physical hardware and software - at cloud scale?

    AWS Hardware Engineering is looking for a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale. You will develop automation, analyze hardware telemetry across tens of thousands of hosts, and build tooling that directly determines whether capacity is available to customers. Your work spans the full stack - from hardware monitoring interfaces, to health diagnostics up through fleet-wide data pipelines and operational dashboards.

    Key job responsibilities

    Fleet Health & Data Analysis

    • Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact
    • Contribute to predictive failure detection using sensor data, error trending, and log correlation
    • Build and maintain operational dashboards and metrics for platform fleet health.
    • Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets

    Systems Development & Automation

    • Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows
    • Develop diagnostic tools for Linux on ARM and x86 architectures
    • Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems
    • Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability

    Cross-Team Collaboration

    • Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions
    • Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers
    • Participate in sprint-based planning and oncall rotation for platform-level escalations

    A day in the life

    Some days you are deep in system event logs chasing a failure pattern across thousands of hosts; other days you are writing automation that eliminates a manual triage workflow entirely. You work with hardware engineers, firmware teams, datacenter operations, and vendor partners - driving quality and reliability from manufacturing through steady-state operations. Located in Cupertino, Seattle, or Denver, you work with global development teams on servers deployed in datacenters worldwide.

    About the team

    AWS Hardware Engineering designs and delivers next-generation cloud infrastructure - the servers, accelerators, and storage platforms that power AWS. Our team builds custom systems for AI training, inference, and compute workloads at global scale. We are directly responsible for launching and maintaining server hardware in the fleet, working across internal development teams, and design partners.

    We value work-life harmony, inclusive culture, and continuous learning. Even if you do not meet all preferred qualifications listed below, we encourage you to apply.

    Numbers & Facts

    LocationCupertino, CA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Skills

    • ARM (Advanced RISC Machine)unmatched
    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Automation System Developmentunmatched
    • Cloud Computingunmatched
    • Computer Engineeringunmatched
    • Computer Firmwareunmatched
    • Computer Maintenanceunmatched
    • Data Analysisunmatched
    • Data Managementunmatched
    • Debugging Skillsunmatched
    • Environmental Issuesunmatched
    • Failure Analysisunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Hardware Designunmatched
    • Hardware Quality Assuranceunmatched
    • Identify Issuesunmatched
    • Javaunmatched
    • Linux Operating Systemunmatched
    • Machine Toolunmatched
    • Manufacturingunmatched
    • Metricsunmatched
    • Microprocessor Architectureunmatched
    • National Intelligence Council (NIC)unmatched
    • Network Operations Centerunmatched
    • On Callunmatched
    • Operating Systemsunmatched
    • PCI Express (PCI-E)unmatched
    • Pattern Analysisunmatched
    • Problem Solving Skillsunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • Server Hardwareunmatched
    • Server Programming/Applicationsunmatched
    • Sprint Planningunmatched
    • Supply Chainunmatched
    • Systems Engineeringunmatched
    • Team Playerunmatched
    • Technical/Engineering Designunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Vehicle Fleetsunmatched
    • x86 Processorsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder