Amazon.com Inc logo

Software Development Manager (EC2 Nitro), EC2 Core Provisioning

Amazon.com Inc

  • Seattle, WA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Automotive Repair and Maintenanceunmatched
    • Best Practicesunmatched
    • Computer Engineeringunmatched
    • Computer Firmwareunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Debugging Skillsunmatched
    • Distributed Computingunmatched
    • Hardware Componentsunmatched
    • Large-Scale Systemsunmatched
    • Machine Learningunmatched
    • Mentoringunmatched
    • Multiplatform/Cross-Platformunmatched
    • People Managementunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Developmentunmatched
    • Software Engineeringunmatched
    • Team Lead/Managerunmatched
    • Team Playerunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Vehicle Fleetsunmatched

    Description

    build, deploy, and manage applications with unparalleled flexibility and efficiency.

    Join our dynamic team, where we apply agentic and machine-learning solutions to one of the hardest problems in the fleet: returning broken servers to production when there is no deterministic signal of what is wrong. You will build a learning, agent-driven decision engine on top of fleet telemetry and repair history, and you will ship it as a production software service that operates across millions of servers in every region, serving every EC2 line of business from core servers to accelerators and UltraServers. Sentinel is a direct lever on unsellable capacity and on the cost of running the fleet, and we are evolving it from human-authored decision rules into a system that recovers capacity on its own.

    We are looking for an experienced Software Development Manager (SDM) to lead this team. The ideal candidate has led teams, thoroughly understands the design, development, and debugging of large-scale distributed systems, and is excited to apply ML and agentic techniques to real hardware-recovery problems. In this role, the manager will work with a broad group of technical teams across hardware, firmware, vetting, and provisioning.

    Key job responsibilities

    • Lead and inspire a team of engineers, providing guidance, mentorship, and support to foster their professional growth.
    • Own the recovery decision engine that returns broken servers to sellable capacity, driving down unsellable rate and the time a host stays stuck. Take on the failures that have no deterministic signal, and evolve the engine from static, human-authored signatures into an agentic, ML-driven system that infers the right repair from fleet outcomes and improves with every recovery.
    • Build and operate this as a production software service - reliable, secure, and observable - running across millions of servers in every region, not a set of offline models or scripts.
    • Debug complex, system-level, multi-component failures across hardware, firmware, BMC, and the provisioning and vetting stack, and turn that diagnosis into automated, repeatable recovery.
    • Collaborate with hardware engineering, firmware, component owners, vetting, and provisioning teams to expand recovery coverage across platforms and drive failures upstream to their root cause so they stop recurring.
    • Raise the bar on the safety of autonomous action on production-bound capacity, holding a high security and operational standard for a service that runs across all regions, including restricted environments.
    • Champion best practices in software engineering, including code quality, testing, automation, and continuous integration and delivery (CI/CD).

    Numbers & Facts

    LocationSeattle, WA
    IndustryRetail
    Company Size10,000 employees or more
    Year Founded1994
    Websitehttp://Amazon.com/militaryroles

    About Company

    At Amazon, we don’t wait for the next big idea to present itself. We envision the shape of impossible things and then we boldly make them reality. So far, this mindset has helped us achieve some incredible things. Let’s build new systems, challenge the status quo, and design the world we want to live in. We believe the work you do here will be the best work of your life.

    Wherever you are in your career exploration, Amazon likely has an opportunity for you. Our research scientists and engineers shape the future of natural language understanding with Alexa. Fulfillment center associates around the globe send customer orders from our warehouses to doorsteps. Product managers set feature requirements, strategy, and marketing messages for brand new customer experiences. And as we grow, we’ll add jobs that haven’t been invented yet.

    It’s Always Day 1
    At Amazon, it’s always “Day 1.” Now, what does this mean and why does it matter? It means that our approach remains the same as it was on Amazon’s very first day – to make smart, fast decisions, stay nimble, invent, and stay focused on delighting our customers. In our 2016 shareholder letter, Amazon CEO Jeff Bezos shared his thoughts on how to keep up a Day 1 company mindset. “Staying in Day 1 requires you to experiment patiently, accept failures, plant seeds, protect saplings, and double down when you see customer delight,” he wrote. “A customer-obsessed culture best creates the conditions where all of that can happen.” You can read the full letter here

    Our Leadership Principles
    Our Leadership Principles help us keep a Day 1 mentality. They aren’t just a pretty inspirational wall hanging. Amazonians use them, every day, whether they’re discussing ideas for new projects, deciding on the best solution for a customer’s problem, or interviewing candidates. To read through our Leadership Principles from Customer Obsession to Bias for Action, visit https://www.amazon.jobs/principles

    Similar Jobs