Product Reliability Engineering Lead

Pyramid, Inc

  • Houston, TX
  • 2 days ago
  • $70–$80 Per Hour
  • Full-time
Want to know if you’re a fit?
Upload your resume and let our AI show you.

Skills

  • AWS Lambdaunmatched
  • Acquisition Strategyunmatched
  • Agile Programming Methodologiesunmatched
  • Amazon Web Services (AWS)unmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Budgetingunmatched
  • Channel Strategiesunmatched
  • Circuit Breakersunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Consultingunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Cross-Functionalunmatched
  • Distribution Channelunmatched
  • Employee Benefitsunmatched
  • Failoverunmatched
  • Failure Analysisunmatched
  • Financial Modelingunmatched
  • Financial Servicesunmatched
  • Functional Testingunmatched
  • Health Insuranceunmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Injectionsunmatched
  • Insuranceunmatched
  • Leadershipunmatched
  • Load Testingunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Multiplatform/Cross-Platformunmatched
  • Offshoringunmatched
  • Operational Improvementunmatched
  • Problem Solving Skillsunmatched
  • Product Engineeringunmatched
  • Radiographyunmatched
  • Reliability Engineeringunmatched
  • Reliability Testingunmatched
  • Reporting Dashboardsunmatched
  • Riskunmatched
  • Short Messaging Service (SMS)unmatched
  • Stress Testingunmatched
  • Team Lead/Managerunmatched
  • Telemetryunmatched
  • Test Plan/Scheduleunmatched
  • Testingunmatched
  • Time Managementunmatched
  • Validation Testingunmatched

Description

Immediate need for a talented Product Reliability Engineering Lead. This is a 12+ Months Contract opportunity with long-term potential and is located in US (Remote). Please review the job description below and contact me ASAP if you are interested. Job ID:26-14182 Pay Range: $70 - $80/hour. Employee benefits include, but are not limited to, health insurance (medical, dental, vision). Key Responsibilities: The Platform Reliability Engineering Lead for the Acquisition Platform is a pivotal role responsible for ensuring the platform scales reliably and sustainably as demand grows across products, channels, and distribution partners. This role shifts reliability left into design and development, embeds observability as a first class platform capability, and leverages AI assisted techniques to accelerate detection, diagnosis, and resolution of issues before they impact partners or customers. The ideal candidate brings deep experience in site reliability engineering, platform observability, and resilience validation, along with the ability to lead teams in treating reliability as a continuous, measurable product discipline rather than a reactive operations function. Reliability Strategy and Architecture Define and lead the reliability strategy for the Acquisition Platform, ensuring alignment with product, platform, and enterprise goals. Establish SLOs, SLIs, and error budgets that tie reliability targets to business outcomes and partner expectations. Shift reliability requirements into early design and development phases so resiliency, failover, and graceful degradation are architected in, not bolted on. Design reliability patterns across platform services, APIs, workflows, and dependent systems both internal and external to client. Observability and Operational Enablement Architect end to end observability across the platform including metrics, structured logging, distributed tracing, and alerting. Establish monitoring standards and dashboards that provide real time visibility into platform health, partner facing services, and integration dependencies. Embed observability into platform services from design through deployment so teams can detect, diagnose, and resolve issues rapidly. Drive adoption of synthetic monitoring and canary deployments to validate production behavior proactively. AI Assisted Reliability Engineering Direct teams in using prompt driven and agent based approaches to automate toil, reduce mean time to recovery, and improve operational consistency. Explore and introduce AI enabled patterns for predictive alerting, automated remediation, and intelligent escalation where appropriate. Apply responsible AI practices with attention to security, data exposure, and operational risk. Reliability Testing and Validation Partner with the Platform Development Engineer in Test to align functional, non functional, and reliability test coverage. Embed reliability validation into CI/CD pipelines so resiliency is continuously tested, not assumed. Validate platform behavior under failure conditions across APIs, integrations, workflows, and experience layers. Collaboration and Stakeholder Engagement Collaborate closely with the Acquisition delivery team and stakeholders to align outcomes with the reliability strategy. Partner with AMS, infrastructure, and other tech teams to ensure clear ownership boundaries and smooth operational handoffs. Key Requirements and Technology Experience: SRE principles – SLOs, SLIs, error budgets, toil reduction, blameless postmortems Observability design – distributed tracing, APM telemetry, structured logging, real time alerting, synthetic monitoring Resilience and fault tolerance – circuit breakers, bulkheads, retry/backoff, graceful degradation, failover validation Chaos engineering and reliability testing – fault injection, load/stress testing, failure mode analysis CI/CD reliability integration – automated reliability gates, canary deployments, feature flags, progressive rollouts AI assisted reliability techniques – anomaly detection, predictive alerting, prompt driven runbook automation, agent based remediation Responsible AI use – including consideration of security, data exposure, and operational risk Cloud native operations – containerized platforms, event driven architectures, infrastructure as code Growth oriented mindset – ability to think beyond constraints of today and identify what is required to build the future Excellent communication skills – ability to translate reliability concerns between engineering, product, and business teams 5+ years of experience in site reliability engineering, platform engineering, or production operations roles Experience defining and operating SLO/SLI frameworks tied to business outcomes Hands on experience designing observability for distributed, API driven platforms Experience with reliability and resiliency testing including chaos engineering and fault injection Experience guiding and mentoring engineers on reliability practices Enterprise scale delivery experience with both onshore and offshore cross functional teams Direct experience applying Agile methodologies in product centric delivery models Experience in financial services, insurance, or retirement services industry AWS operational experience – CloudWatch, X Ray, Fault Injection Simulator, ECS/EKS, Lambda, EventBridge Experience integrating reliability practices with DevSecOps and CI/CD pipelines Familiarity with AI/ML driven operations tools and incident management platforms. Our client is a leading Life Insurance/Retirement Solutions Industry and we are currently interviewing to fill this and other similar contract positions. If you are interested in this position, please apply online for immediate consideration. Pyramid Consulting, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. By applying to our jobs you agree to receive calls, AI-generated calls, text messages, or emails from Pyramid Consulting, Inc. and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here.

Numbers & Facts

LocationHouston, TX
Job TypeFull-time
Salary$70–$80 Per Hour

Similar Jobs

See more jobs