With a career at The Home Depot, you can be yourself and also be part of something bigger.
Position Purpose:
As Manager, Reliability Engineering for Transportation and Delivery Fulfillment, you lead the team that runs and improves the supply chain systems that move freight to our stores and distribution centers and get product to our customers. Your team spans onshore and offshore engineers, including staff engineers, and you hold operational and engineering work to a single set of goals. You own these systems 24x7. When a major incident is declared, you are the escalation point: you open the call, direct the technical response, and tell business and executive stakeholders about the impact, the scope, and the expected recovery. Those updates go out on a cadence you publish in advance. You run both domains as one practice, with one on-call and escalation model, one problem review, and one prioritized reliability backlog. You map the Critical User Journeys (CUJs) the business depends on, then set and enforce Service Level Objectives (SLOs) for availability and performance against them. Problem management, automation, and applied AI are how you remove recurring operational work for good. You keep changes from causing incidents, and you deliver resilience testing and security remediation on the dates we commit to. You hire, develop, and recognize your engineers.
Key Responsibilities:
30% Delivery & Execution:
10% Support & Enablement:
50% People:
10% Learning:
Direct Manager/Direct Reports:
Travel Requirements:
Physical Requirements:
Working Conditions:
Minimum Qualifications:
Preferred Qualifications:
Experience leading combined software engineering and day-to-day operational work with accountability for application development, reliability, and operational excellence.
Experience managing globally distributed engineering teams, including onshore and offshore resources, across multiple application domains.
Experience owning 24x7 on-call, incident response, and escalation processes for business-critical systems, including driving post-incident reviews and corrective actions.
Experience establishing problem management practices focused on root cause analysis, recurring issue elimination, and continuous service improvement.
Experience implementing automation and AI-driven solutions to reduce manual operational effort, improve efficiency, and accelerate incident resolution.
Experience supporting supply chain, transportation, fulfillment, logistics, or customer delivery platforms in a large-scale enterprise environment.
Strong executive communication skills with the ability to translate complex technical issues into clear business impact, risk, and recovery plans.
Experience defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and service performance metrics across complex application ecosystems.
Proven ability to recruit, develop, mentor, and retain high-performing engineering talent while fostering a culture of accountability and continuous learning.
Experience developing technology roadmaps, driving quarterly planning, and aligning engineering priorities with business objectives.
Experience managing engineering capacity, operational workload, cloud consumption, and technology budgets in a cost-conscious environment.
Experience overseeing modern CI/CD pipelines, release management processes, and change governance practices to support reliable software delivery.
Experience driving system resiliency, disaster recovery readiness, capacity planning, and performance optimization for high-volume platforms.
Strong knowledge of observability platforms, including logging, metrics, tracing, and monitoring tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, or Elastic.
Experience operating cloud-native applications on Google Cloud Platform, or Microsoft Azure, including container platforms such as Kubernetes.
Experience with Infrastructure as Code and configuration management tools such as Terraform, Ansible, Chef, or Puppet.
Experience supporting hybrid technology environments that include cloud services, on-premises platforms, vendor-supported applications, and relational databases.
Proficiency in at least one modern programming or scripting language such as Python, Java, Go, or Bash.
Experience partnering with software engineering, quality engineering, security, infrastructure, and business teams to improve application reliability and operational effectiveness.
Experience collaborating with security and compliance teams to meet regulatory and corporate policy requirements, including PCI-DSS and SOC 2.
Working knowledge of identity, access management, secrets management, credential lifecycle management, and security best practices in enterprise environments.
Minimum Education:
Preferred Education:
Minimum Years of Work Experience:
Preferred Years of Work Experience:
Minimum Leadership Experience:
Preferred Leadership Experience:
Certifications:
Competencies:
For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:
| Location | Georgia (Remote) |
| Industry | Construction - Residential & Commercial/Office |
| Company Size | 10,000 employees or more |
| Website | https://careers.homedepot.com/ |
As an essential retailer to the communities we serve, our stores are open and our warehouse distribution centers are running. We’re taking measures to keep our associates safe, including:
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder