DevOps Engineer IV- 4P/702

4P Consulting
  • Atlanta, Georgia
    30+ days ago

    Job Description

     DevOps Engineer IV – Site Reliability Engineer / Observability

    Location: Atlanta, Ga                                                                                                                                                 Contract : 3Years
    Experience: Senior Level

    Client- Southern Company Services.

    Job Summary

    We are seeking an experienced DevOps Engineer IV / Site Reliability Engineer (SRE) with strong hands-on experience in observability, telemetry, monitoring, and service reliability. The ideal candidate will have deep knowledge of Grafana, OpenTelemetry (OTEL), PromQL, and application/system instrumentation.

    This role will partner with engineering, operations, and application teams to improve service reliability, telemetry quality, alerting maturity, and operational visibility across complex environments.

    Key Responsibilities

    • Design, implement, and support monitoring and observability solutions.
    • Build dashboards, alerts, and telemetry solutions using Grafana and related tools.
    • Implement OpenTelemetry standards for application and system instrumentation.
    • Write and optimize PromQL queries for monitoring and reliability insights.
    • Improve alerting quality, reduce noise, and create actionable alerts.
    • Troubleshoot application and infrastructure issues using logs, metrics, and traces.
    • Support incident response, root cause analysis, and reliability improvements.
    • Collaborate with engineering, operations, and application teams.

    Required Qualifications

    • Strong experience as a DevOps Engineer, SRE, Observability Engineer, or similar role.
    • Hands-on experience with Grafana, OpenTelemetry, and PromQL.
    • Experience with application and system instrumentation.
    • Strong understanding of logs, metrics, traces, alerting, and service reliability.
    • Ability to design monitoring solutions across complex environments.
    • Strong troubleshooting, analytical, communication, and collaboration skills.

    Preferred Qualifications

    • Experience with Prometheus, Loki, Tempo, Kubernetes, containers, cloud platforms, or microservices.
    • Familiarity with CI/CD, automation, infrastructure-as-code, incident response, SLIs, SLOs, and reliability metrics.

    Key Skills

    DevOps, SRE, Observability, Grafana, OpenTelemetry, OTEL, PromQL, Prometheus, Monitoring, Alerting, Logs, Metrics, Traces, Instrumentation, Incident Response, Root Cause Analysis.

    Numbers & Facts

    LocationAtlanta, Georgia

    Skills

    • Analysis Skillsunmatched
    • Automationunmatched
    • Cloud Computingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • DevOpsunmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • Instrumentationunmatched
    • Metricsunmatched
    • Microservicesunmatched
    • Quality Managementunmatched
    • Query Optimizationunmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • Team Playerunmatched
    • Telemetryunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder