Position Title: Senior Site Reliability Engineer (SRE) - Mobile & Digital Observability Location: Atlanta, GA(Onsite) Duration: Long Term Contract
Job Description:
The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems. This role is part of the Observability team and works closely with mobile, web, and backend engineering teams to ensure full visibility into customer journeys and user experience. The focus is on building and operating Real User Monitoring (RUM), synthetic monitoring, and end-to-end telemetry correlation using Splunk and SignalFx, ensuring issues are detected before customer impact. This is not a monitoring-only role—it requires active involvement in instrumentation, release observability, and reliability engineering.
Required Skills:
Strong experience with Splunk (logs, dashboards) and SignalFx (metrics/APM)
Hands-on with RUM and synthetic monitoring tools
Experience with mobile observability (iOS/Android), including performance monitoring and crash analysis
Strong understanding of distributed systems and microservices (Java, Node.js)
Experience with Azure (AKS, App Services, APIM)
Ability to correlate frontend issues with backend services
Experience with CI/CD pipelines and observability in release processes
Role and Responsibilities:
Define and enforce SLIs, SLOs, and error budgets for critical customer journeys (ordering, checkout, payments)
Own end-to-end observability across Mobile, Web, and POS platforms
Implement and operate RUM and synthetic monitoring for customer-facing journeys
Build mobile-first monitoring coverage including app performance, crash rates, API performance, and user journey tracking
Use Splunk and SignalFx to design dashboards, detectors, and actionable alerts
Enable correlation across mobile, CDN, API, backend systems using logs, metrics, and traces
Partner with engineering teams for instrumentation, SDK integration, and embedding observability into releases
Analyze telemetry to detect post-release issues, device/OS-specific failures, and network degradation
Lead response for P1/P2 incidents and drive root cause analysis
Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.
We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.
Skills
Analysis Skillsunmatched
Androidunmatched
Application Programming Interface (API)unmatched
Budgetingunmatched
Content Delivery Network (CDN)unmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Customer Relationsunmatched
Distributed Computingunmatched
Instrumentationunmatched
Instrumentation Engineeringunmatched
Javaunmatched
Metricsunmatched
Microservicesunmatched
Microsoft Windows Azureunmatched
Node.jsunmatched
Operating Systemsunmatched
Operational Improvementunmatched
Performance Analysisunmatched
Point of Sale (POS) Systemsunmatched
Quality System Requirements (QSR)unmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Root Cause Analysisunmatched
Splunkunmatched
Telemetryunmatched
User Interface/Experience (UI/UX)unmatched
Web Programmingunmatched
iOSunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.