We're seeking a future team member for the role of Sr. Site Reliability Automation Engineer to join our Technology team. This role is located in Lake Mary, FL and Pittsburgh, PA.
In this role, you'll make an impact in the following ways:
Design and implement end-to-end observability (logs, metrics, traces) across distributed systemsBuild Observability & Monitoring
Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
Identify gaps in monitoring and drive adoption of best practices
Drive Automation & Reduce Toil
Identify repetitive operational work and automate it using code and tooling
Build self-healing and auto-remediation solutions
Enable scalable, reliable processes through automation and engineering rigor
Improve operational efficiency across production environments
Support Production & Incident Triage
Troubleshoot and resolve complex production issues across distributed systems
Participate in incident management, triage, and root cause analysis
Improve monitoring and automation based on recurring incident patterns
Collaborate with support and engineering teams to improve system stability
Improve Reliability & Performance
Define and measure service health using SLIs/SLOs and key performance metrics
Identify system bottlenecks and reliability risks
Contribute to performance optimization and capacity planning
Provide input into system architecture to improve resilience and scalability
To be successful in this role, we're seeking the following:
3-6 years of experience in Site Reliability Engineering, Software Engineering
Strong programming background in Java (preferred) or another modern language
Experience with at least one observability platform:
•AppDynamics, Dynatrace, Grafana, or Splunk
Hands-on experience supporting and troubleshooting production systems
Strong analytical and problem-solving skills
Ability to identify inefficiencies and drive automation
Preferred Qualifications
Experience with distributed systems or microservices architectures
Familiarity with CI/CD pipelines and DevOps practices
Exposure to cloud platforms and/or Kubernetes
Experience scripting (Python, Bash, etc.) for automation
Knowledge of SRE concepts like observability, incident management, and reliability engineering
Numbers & Facts
Location
Pittsburgh, PA
Skills
Analysis Skillsunmatched
Automationunmatched
Automation Engineeringunmatched
Bash Scriptingunmatched
Best Practicesunmatched
Capacity and Performance Managementunmatched
Cloud Computingunmatched
Computer Programmingunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
DevOpsunmatched
Distributed Computingunmatched
Identify Issuesunmatched
Incident Managementunmatched
Javaunmatched
Machine Toolunmatched
Metricsunmatched
Microservicesunmatched
Operational Improvementunmatched
Operational Strategyunmatched
Performance Managementunmatched
Performance Metricsunmatched
Performance Tuning/Optimizationunmatched
Problem Solving Skillsunmatched
Production Supportunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Root Cause Analysisunmatched
Scripting (Scripting Languages)unmatched
Software Engineeringunmatched
Splunkunmatched
System Architectureunmatched
Systems Engineeringunmatched
Systems Reliabilityunmatched
Telemetryunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.