Senior Site Reliability Engineer

O.C. Tanner

  • Salt Lake City, UT
  • 5 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Web Services (AWS)unmatched
    • Apache ActiveMQunmatched
    • Apache Kafkaunmatched
    • Automationunmatched
    • Automation Systemsunmatched
    • Best Practicesunmatched
    • Business Servicesunmatched
    • Cloud Applicationsunmatched
    • Cloud Computingunmatched
    • Continuous Improvementunmatched
    • Customer Experienceunmatched
    • DevOpsunmatched
    • Health Maintenanceunmatched
    • High Availabilityunmatched
    • Home Automationunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Javaunmatched
    • Machine Toolunmatched
    • Metricsunmatched
    • Operational Improvementunmatched
    • Performance Managementunmatched
    • PostgreSQLunmatched
    • Problem Solving Skillsunmatched
    • Product Supportunmatched
    • Production Systemsunmatched
    • Programming Languagesunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Managementunmatched
    • Redisunmatched
    • Reliability Engineeringunmatched
    • Risk Managementunmatched
    • Root Cause Analysisunmatched
    • Simple Queue Service (SQS)unmatched
    • Software Development Lifecycle (SDLC)unmatched
    • Software Engineeringunmatched
    • Software Testingunmatched
    • Software as a Service (SaaS)unmatched
    • Systems Scalabilityunmatched

    Description

    As a Senior Site Reliability Engineer, you will help define the future of reliability for our world‑class employee recognition platform. You'll leverage software engineering, automation, and cloud‑native technologies to build and operate highly available, scalable systems that serve millions of users. We're looking for someone who is passionate about reliability engineering, continuous improvement, and building self‑healing platforms that enable development teams to move faster while delivering exceptional customer experiences.Key ResponsibilitiesImprove the availability, scalability, and performance of cloud‑native applications through automation, monitoring, and engineering best practices.Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service‑level objectives (SLOs) that enable proactive issue detection and resolution.Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues.Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle.Champion a reliability‑first engineering culture by establishing automation standards, monitoring best practices, shift‑left quality approaches and shared ownership models that proactively improve resilience, reduce operational risk, and protect the availability of business‑critical services.Collaborate with global engineering teams in a follow‑the‑sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services.Participate in an on‑call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks.Required Qualifications5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence.Experience operating large‑scale, customer‑facing SaaS platforms with high availability and uptime requirements.Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability‑focused engineering solutions.Deep experience with modern Infrastructure‑as‑Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes.Hands‑on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms.Strong knowledge of AWS services and Kubernetes in production environments.Deep understanding of monitoring, logging, and distributed tracing for complex systems.Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle.Comfortable with participating in on‑call rotations and handling high‑pressure environments.Bonus QualificationsExperience with multiple cloud or cloud‑agnostic environments.Familiarity with security, compliance, and governance frameworksExperience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event‑driven technologies.#J-18808-Ljbffr

    Numbers & Facts

    LocationSalt Lake City, UT

    Similar Jobs

    See more jobs