Site Reliability Engineer - New York (Remote)

Georgia Tek Systems
  • NY, NY
  • Remote
    7 days ago

    Job Description

    Job Title: Site Reliability Engineer (SRE) – Dynatrace
    Location: New York (Remote)
    Experience: 6–10 Years
    Rate: DOE
    Job Description
    We are looking for a highly skilled Site Reliability Engineer (SRE) with strong expertise in Dynatrace monitoring and observability solutions. The ideal candidate will be responsible for ensuring the reliability, scalability, and performance of enterprise applications and infrastructure across cloud and on-prem environments.
    The candidate should have hands-on experience with monitoring, automation, troubleshooting, cloud platforms, and modern DevOps practices.
    Key Responsibilities
    • Design, implement, and maintain end-to-end monitoring solutions using Dynatrace.
    • Configure dashboards, alerts, problem detection rules, and observability frameworks.
    • Monitor application performance, infrastructure health, and distributed systems.
    • Troubleshoot production issues and perform root cause analysis to improve system reliability.
    • Work closely with DevOps, Cloud, and Application teams to optimize system performance.
    • Automate operational tasks using scripting languages such as Python, Bash, or Shell.
    • Support and manage containerized environments using Docker and Kubernetes.
    • Implement and maintain CI/CD pipelines using tools like Jenkins, GitLab CI/CD, or Azure DevOps.
    • Ensure high availability, scalability, and resiliency of systems and services.
    • Participate in incident response, on-call rotations, and performance tuning activities.
    • Create and maintain technical documentation, runbooks, and operational procedures.
    Required Skills & Qualifications
    • 6–10 years of experience in Site Reliability Engineering, DevOps, or Production Support roles.
    • Strong hands-on expertise with Dynatrace including monitoring, alerting, dashboards, and problem analysis.
    • Solid understanding of observability, logging, monitoring frameworks, and APM tools.
    • Experience working with cloud platforms such as AWS, Azure, or GCP.
    • Strong knowledge of Linux/Unix administration and troubleshooting.
    • Experience with Docker, Kubernetes, and container orchestration.
    • Hands-on experience with CI/CD tools including Jenkins, GitLab, or Azure DevOps.
    • Strong scripting and automation skills using Python, Bash, or Shell scripting.
    • Good understanding of microservices architecture and distributed systems.
    • Experience with incident management, root cause analysis, and system performance optimization.
    • Excellent communication and problem-solving skills.
    Preferred Qualifications
    • Experience with Infrastructure as Code tools such as Terraform or Ansible.
    • Exposure to logging tools like Splunk, ELK Stack, or Grafana.
    • Knowledge of Agile/Scrum methodologies.
    • Relevant certifications in Cloud, Kubernetes, or Dynatrace are a plus.

    Numbers & Facts

    LocationNY, NY (
    Remote
    )

    Skills

    • Agile Programming Methodologiesunmatched
    • Amazon Web Services (AWS)unmatched
    • Ansibleunmatched
    • Bash Scriptingunmatched
    • Cloud Applicationsunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • DevOpsunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • Enterprise Applicationsunmatched
    • GCP (Good Clinical Practices)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Jenkinsunmatched
    • Linux Administrationunmatched
    • Microservicesunmatched
    • Microsoft Windows Azureunmatched
    • On Callunmatched
    • Operations Processesunmatched
    • Performance Analysisunmatched
    • Performance Tuning/Optimizationunmatched
    • Problem Solving Skillsunmatched
    • Production Supportunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • Scripting (Scripting Languages)unmatched
    • Scrum Project Management and Software Developmentunmatched
    • Splunkunmatched
    • Systems Reliabilityunmatched
    • Technical Writingunmatched
    • United States Department of Energy (DOE)unmatched
    • Unix Shell Programmingunmatched
    • Unix System Administrationunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder