SENIOR SITE RELIABILITY ENGINEER (SRE)

Compunnel
  • Atlanta, GA
  • $75–$80 Per Hour
  • Full-time
  • Instant Apply
3 days ago

Job Description

We are seeking a Senior Site Reliability Engineer with a strong web and mobile application development background to drive observability, reliability, and performance across customer-facing digital platforms. The ideal candidate is a software engineer who applies an SRE mindset to application design, instrumentation, automation, production readiness, and incident prevention.

This role is focused on digital experience reliability across web, mobile, APIs, and critical customer journeys. The successful candidate will help establish proactive monitoring capabilities so that user-impacting issues can be detected through telemetry and synthetic validation before they are reported by customers.

 

Job Responsibilities

Web & Mobile Application Reliability

Partner with web, mobile, API, platform, product, and quality engineering teams to improve the reliability of customer-facing applications.

Understand browser, frontend, mobile, network, API, and third-party dependency behavior to troubleshoot customer experience issues.

Review application architecture and code to identify performance, reliability, and observability gaps.

Contribute code and engineering changes that improve instrumentation, fault tolerance, diagnostics, and operational readiness.

Evaluate application behavior across device types, operating systems, browser versions, releases, and network conditions.

 

Digital Experience Observability

Design and implement end-to-end observability across web applications, iOS and Android applications, APIs, cloud services, and external dependencies.

Implement and improve Real User Monitoring (RUM), mobile application monitoring, Application Performance Monitoring (APM), logging, metrics, and distributed tracing.

Instrument applications using OpenTelemetry or equivalent standards and enable correlation from the user session through downstream services.

Monitor frontend performance and user experience signals, including page and screen performance, errors, crashes, failed transactions, Core Web Vitals, and API behavior.

Create actionable dashboards and alerts centered on customer and business impact rather than infrastructure health alone.

Use session replay and digital analytics data, where available, to support diagnosis of user-impacting issues.

 

Synthetic Monitoring & Customer Journey Validation

Design, develop, and maintain browser, API, and mobile synthetic tests for business-critical customer journeys.

Build proactive validation for flows such as sign-in, account creation, browsing, ordering, checkout, payment, loyalty, and account management.

Develop reusable test components and integrate synthetic coverage with application changes and release processes.

Ensure synthetic tests provide meaningful diagnostics, telemetry correlation, and actionable alerting when a journey fails or degrades.

Partner with security, engineering, and platform teams to address bot protection, test data, authentication, environment, and third-party dependency constraints.

 

Reliability & Resiliency Engineering

Define and manage SLIs, SLOs, and Error Budgets for critical applications and customer journeys.

Drive observability and reliability requirements into architecture reviews, design decisions, production readiness reviews, and release planning.

Implement and validate reliability patterns such as timeouts, retries, circuit breakers, bulkheads, rate limiting, caching, and graceful degradation.

Participate in performance testing, dependency failure testing, capacity planning, and resiliency validation.

Identify systemic reliability risks and drive corrective engineering actions before customer impact occurs.

 

Software Engineering & Automation

Develop production-quality tools, libraries, services, and automation that improve reliability and reduce operational toil.

Create reusable instrumentation libraries, telemetry standards, and self-service capabilities for engineering teams.

Integrate observability and reliability validation into CI/CD pipelines.

Automate diagnostics, health validation, release verification, incident response, and recovery activities.

Perform code reviews and mentor engineers on application observability and reliability practices.

Incident Management & Continuous Improvement

Provide technical leadership during customer-impacting incidents involving web, mobile, APIs, or dependent services.

Use application code, logs, traces, metrics, RUM, synthetic results, and session data to accelerate diagnosis.

Drive blameless post-incident reviews and ensure corrective actions address systemic causes.

Improve detection, diagnosis, response, recovery, and prevention capabilities through engineering changes.

Track reliability outcomes and use operational learnings to improve application design and monitoring coverage.

Numbers & Facts

LocationAtlanta, GA
Job TypeFull-time
Salary$75–$80 Per Hour
Year Foundednull

Qualifications

Application Development Experience

7+ years of software engineering experience, including recent hands-on development or significant enhancement of customer-facing web and/or mobile applications.

Strong experience with modern web development technologies such as JavaScript, TypeScript, React, Angular, Vue.js, Next.js, or comparable frameworks.

Experience with mobile application development or architecture using iOS/Swift, Android/Kotlin, React Native, Flutter, or comparable technologies.

Experience consuming, troubleshooting, and instrumenting REST or GraphQL APIs from web and mobile applications.

Ability to read, debug, modify, test, and review production application code.

Strong understanding of browser behavior, mobile application lifecycle, frontend performance, network calls, state management, and client-side error handling.

 

Observability & Digital Experience

Hands-on experience implementing RUM, mobile monitoring, synthetic monitoring, APM, distributed tracing, centralized logging, metrics, and alerting.

Experience with OpenTelemetry or comparable application instrumentation standards.

Experience with one or more observability platforms such as Splunk, Datadog, Dynatrace, New Relic, Grafana, AppDynamics, or equivalent.

Experience troubleshooting customer journeys using correlated frontend, mobile, API, backend, and third-party telemetry.

Understanding of Core Web Vitals, frontend performance, mobile crash analytics, release health, and user experience monitoring.

 

SRE, Cloud & Delivery Practices

Strong understanding of SRE principles, including SLIs, SLOs, Error Budgets, production readiness, incident management, and toil reduction.

Experience with distributed systems, microservices, cloud platforms, containers, and Kubernetes.

Experience with Git-based development workflows, automated testing, and CI/CD pipelines.

Working knowledge of Infrastructure as Code and cloud-native deployment practices.

Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.

 

Preferred Qualifications

Experience building observability for high-volume commerce, payments, loyalty, ordering, or other customer-facing digital platforms.

Experience creating browser, API, and mobile synthetic test suites and managing associated test data.

Experience with session replay or digital experience analytics platforms such as Quantum Metric, FullStory, or equivalent.

Experience with CDN, edge, and web security platforms such as Cloudflare, Akamai, or equivalent.

Experience with chaos engineering, fault injection, performance engineering, and resiliency testing.

Experience with accessibility, mobile release monitoring, app-store release health, or device and OS segmentation.

Skills

  • Application Programming Interface (API)unmatched
  • Automationunmatched
  • Customer Relationsunmatched
  • Instrumentationunmatched
  • Mobile Web Programmingunmatched
  • Reliability Engineeringunmatched
  • Software Designunmatched
  • Software Engineeringunmatched
  • Telemetryunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder