Position Objective: In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAV's Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.
Duties and Responsibilities:
Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
Automate cloud operations to reduce manual work and improve consistency.
Collaborate with development and operations teams to enhance deployment and incident response practices.
Implement and refine monitoring, alerting, dashboards, and operational reporting.
Identify reliability risks and recommend improvements to cloud architecture and processes.
Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.
Numbers & Facts
Location
Atlanta, GA
Skills
Cloud Architectureunmatched
Cloud Computingunmatched
Data Analysisunmatched
Identify Issuesunmatched
Incident Responseunmatched
Microsoft Windows Azureunmatched
Operational Improvementunmatched
Problem Solving Skillsunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Risk Analysisunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.