Pay Rate Range: $ 49.24 - 53.03/hr.
Job Title: Platform Reliability Engineer
Job Description:
The Platform Reliability Engineer establishes and operates the shared reliability and operational engineering capabilities that allow EDP's digital platforms to be deployed, observed, diagnosed, recovered, and scaled predictably. The role owns the technical mechanisms for CI/CD, environments, observability, operational readiness, recovery, and production reliability across SAP Commerce Cloud and the MOSAIC composable platform.
This is a capability-building role — you are not inheriting a mature operational system; you are building the practices, automation, controls, and operational capabilities that do not yet exist.
Key Responsibilities
CI/CD and Developer Platform
Own CI/CD pipelines, build and release engineering, and developer experience tooling across EDP-managed digital platforms.
Implement and operate automated pipeline controls and enforcement mechanisms defined by Engineering Excellence, Quality Engineering, Security, and Architecture.
Automate repetitive operational activities, environment provisioning and configuration, deployment controls, diagnostic workflows, and recovery procedures to reduce manual intervention and operational variance.
Environment Engineering
Own Vercel platform configuration, environment strategy, and deployment governance.
Own environment management across the estate — CCv2 and Vercel environments, environment consistency, configuration governance, drift detection, and provisioning.
Observability and SLOs
Define and implement the platform observability architecture, including telemetry standards, service naming, correlation identifiers, distributed trace propagation, log structure, metric conventions, dashboard patterns, retention expectations, and ownership boundaries across platform components and integrations.
Develop actionable alerting and operational runbooks.
Establish the SLI/SLO framework and telemetry, partnering with business, product, and technology owners to define appropriate service objectives for critical services and customer journeys.
Establish synthetic monitoring for critical customer journeys and service dependencies, aligned with the Critical Journey Registry and defined SLOs.
Establish synthetic monitoring for critical customer journeys and service dependencies, aligned with the Critical Journey Registry and defined SLOs.
Reliability and Recovery
Define and maintain operational reliability standards, patterns, and controls, and represent reliability and operability concerns in cross-functional technical governance.
Design and implement repeatable, tested rollback and recovery procedures, replacing manual recovery processes with defined, predictable operations.
Establish performance baselines and monitor for regression across releases. Provide the infrastructure and environment support for performance and load testing in partnership with Quality Engineering.
Establish capacity and scalability baselines for critical platform components; identify emerging constraints and provide evidence for scaling, configuration, and architectural decisions.
Operational Readiness and Security Controls
Define operational readiness criteria — what must be true before a new deployment, SI delivery, or BU onboarding goes live — and apply them as a gate before EDP accepts production operational responsibility.
Implement dependency scanning and security scanning mechanisms in the CI/CD pipeline, operating the enforcement infrastructure that Central Security policy and Engineering Excellence standards require.
Incident and Production Engineering
Lead or coordinate technical diagnosis of complex production incidents that span platform components, vendors, or integration boundaries, while preserving service ownership with the responsible engineering teams.
Lead post-incident reviews and ensure resulting reliability and instrumentation improvements are implemented.
Identify recurring failure patterns and systemic reliability risks from incidents, telemetry, and operational data; drive problem-management actions that eliminate repeat failure rather than repeatedly treating symptoms.
Qualifications
8+ years of relevant experience in platform engineering, infrastructure engineering, DevOps, production engineering, or site reliability engineering, including experience operating at senior or lead level.
Demonstrated experience establishing observability and reliability practices for distributed enterprise applications — building operational capability from low or zero baseline, not solely operating within an already-mature environment.
Hands-on experience with observability architecture and implementation — logging, metrics, distributed tracing, dashboards, alerting, and production monitoring using enterprise observability platforms such as Datadog or Dynatrace.
Strong CI/CD pipeline ownership and build/release engineering experience, including CI pipeline monitoring and optimization using tools such as Datadog CI Visibility or equivalent.
SAP Commerce Cloud (CCv2) operational experience strongly preferred, including familiarity with its deployment model, build process, monitoring, performance characteristics, and platform-specific constraints; comparable experience operating large-scale Java commerce or enterprise application platforms considered.
Vercel or comparable cloud/edge platform experience at production scale.
Experience with environment management, configuration governance, and provisioning across heterogeneous deployment targets.
Strong experience with infrastructure-as-code and configuration-as-code practices for repeatable environment and platform configuration, using Terraform or comparable tooling where supported by the underlying platforms.
Strong scripting and automation capability using languages such as Python, TypeScript/JavaScript, Bash, or PowerShell.
Experience with performance testing infrastructure and environment provision for load testing; familiarity with tools such as k6 or comparable.
Experience defining and operationalizing service level indicators (SLIs), service level objectives (SLOs), and other measures of production health.
Working knowledge of automated testing practices and how unit, integration, contract, and end-to-end testing fit into CI/CD and release controls.
Experience implementing dependency scanning and security scanning in CI/CD pipelines (Snyk, SonarCloud, or equivalent).
Strong incident management, production troubleshooting, and post-incident review experience.
Demonstrated ability to establish reliability, observability, and operational standards across multiple engineering teams.
Strong ability to influence engineering teams, establish operational practices across organizational boundaries, and translate reliability risks into actionable technical and business decisions.
Skills: Digital : DevOps Continuous Integration and Continuous Delivery (CI/CD)
Experience Required: 8-10 years
| Skills: | | Category | Name | Required | Importance | Experience |
|---|
| SkillCategoryTest1_MN | Digital : DevOps Continuous Integration and Continuous Delivery (CI/CD) | Yes | 1 | >7 years | |
|
|---|