Abbott logo

Senior Site Reliability Engineer

Abbott

  • Sunnyvale, California
  • 3 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Analysis Skillsunmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Best Practicesunmatched
    • Biologyunmatched
    • Budgetingunmatched
    • Business Continuity Planning (BCP)unmatched
    • Capacity Managementunmatched
    • Cardiac Monitoringunmatched
    • Cardiologyunmatched
    • Cloud Architectureunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • Cross-Functionalunmatched
    • Customer Relationsunmatched
    • Customer Relationship Management (CRM)unmatched
    • DNS (Domain Name System)unmatched
    • Data Collectionunmatched
    • Debugging Skillsunmatched
    • Design Patterns Programming Methodologiesunmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • Documentationunmatched
    • EEO Regulationsunmatched
    • English Lawunmatched
    • Environmental Sciencesunmatched
    • Follow Throughunmatched
    • Go Programming Language (Golang)unmatched
    • HIPAA (Health Insurance Portability and Accountability Act)unmatched
    • HTTP (HyperText Transport Protocol)unmatched
    • Healthcareunmatched
    • High Availabilityunmatched
    • High Throughputunmatched
    • Home Automationunmatched
    • Identify Issuesunmatched
    • Implantsunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • International Healthunmatched
    • Linux Operating Systemunmatched
    • Load Balancingunmatched
    • Machine Toolunmatched
    • Marketingunmatched
    • Medical Diagnosisunmatched
    • Medical Equipmentunmatched
    • Medical Productsunmatched
    • Messaging Middlewareunmatched
    • Microservicesunmatched
    • Microsoft Windows Azureunmatched
    • On Callunmatched
    • Operations Processesunmatched
    • Patient Careunmatched
    • Patient Safetyunmatched
    • Problem Solving Skillsunmatched
    • Product Developmentunmatched
    • Production Systemsunmatched
    • Programming Languagesunmatched
    • Python Programming/Scripting Languageunmatched
    • RMONunmatched
    • Regulationsunmatched
    • Reliability Engineeringunmatched
    • Resource Managementunmatched
    • Root Cause Analysisunmatched
    • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
    • Safety Trainingunmatched
    • Software Engineeringunmatched
    • Strategic Planningunmatched
    • Surveillanceunmatched
    • Systems Engineeringunmatched
    • Systems/Internals Programmingunmatched
    • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
    • Testingunmatched
    • Windows PowerShellunmatched

    Description

    Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.

         

    JOB DESCRIPTION:

    About the Role

    This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.

    We are seeking a highly skilled and mission-driven Senior Site Reliability Engineer (SRE) to join our DevOps team. In this critical role, you will be responsible for ensuring the reliability, scalability, performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists, and care teams automatically collect and review data from patients with implanted cardiac devices.

    This isn't just about keeping servers up; it's about building and maintaining the resilient backbone for systems where failure is not an option, and where our success directly impacts patient care around the world. You will embed within our DevOps team, acting as a bridge between development and operations.

    What You'll Do

    Design, implement, and maintain highly available, fault-tolerant, and resilient systems that meet demanding uptime and safety requirements. Identify and eliminate performance bottlenecks in software and infrastructure, ensuring low-latency, high-throughput, and real-time responsiveness for customer-facing services. Define, monitor, and uphold Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Develop and implement comprehensive monitoring, logging, tracing, and alerting solutions to provide deep insights into system health and behavior at scale. Automate away manual operational tasks, from provisioning and deployment to testing and recovery. Develop and implement strategies for scaling our services and infrastructure to meet evolving business demands, including distributed systems and cloud deployments in Azure. Work closely with software engineering, security, quality, and compliance teams to integrate SRE best practices into our operational processes and infrastructure, ensuring the integrity, availability, and confidentiality of our systems that carry sensitive patient health data. Create clear, concise, and comprehensive documentation, runbooks, and playbooks for operational procedures. Lead blameless postmortem processes following incidents and drive systematic follow-through on action items. Work with a multi-disciplinary team on challenging problems in a fast-paced environment, contributing across architecture reviews, incident response, capacity planning, and reliability roadmap planning.

    Required Qualifications

    • Bachelor's in Computer Science, Software Engineering, Systems Engineering, or a related technical discipline; equivalent professional experience will be considered.
    • Minimum 7 years of experience working in site reliability, software Engineering , Systems Engineering or a related technical discipline
    • Excellent communication skills with the demonstrated ability to work effectively in cross-functional teams, translate technical complexity for non-technical stakeholders, and collaborate with development, quality, security, marketing, and regulatory teams.
    • Strong analytical, problem-solving, and debugging skills with a methodical and structured approach to diagnosing complex, distributed system issues under pressure — including production incidents with patient safety implications.
    • Proficiency in at least one systems or automation programming language (e.g., Python, Go, Bash, PowerShell) for building tooling, automation, and operational systems.
    • Demonstrated expertise with Microsoft Azure — including Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, and related managed services.
    • Container orchestration expertise — hands-on production experience with Kubernetes and Docker at scale, including deployment strategies, resource management, and cluster operations.
    • Observability platform experience with tools such as Prometheus, Grafana, the ELK/EFK stack, Datadog, Azure Monitor, or similar enterprise-grade monitoring and tracing platforms.
    • Experience designing and operating CI/CD pipelines for continuous delivery of software in production environments, including safe deployment strategies such as blue/green, canary, and feature flag-gated rollouts.
    • Deep understanding of distributed systems — including load balancing, service meshes, microservices, message queues, and fault-tolerant design patterns.
    • Solid Linux & networking fundamentals — DNS, TCP/IP, HTTP/S, TLS, load balancing, and networking in cloud environments.
    • Incident management experience — including on-call rotation participation, structured incident response, root cause analysis (RCA), and systematic prevention of recurrence.

    Preferred Qualifications

    • Experience working in a regulated healthcare, medical device, or life sciences environment, with familiarity in compliance frameworks such as HIPAA.
    • Relevant professional certifications such as: Microsoft Certified Azure DevOps Engineer Expert, Azure Solutions Architect Expert, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Security Specialist (CKS).
    • Background in cost optimization for cloud-native architectures, including FinOps practices for Azure environments.

    Experience contributing to or driving disaster recovery (DR) design, business continuity planning, and tabletop exercises.

         

    The base pay for this position is

    $90,000.00 – $180,000.00

    In specific locations, the pay range may vary from the range posted.

         

    JOB FAMILY:

    Product Development

         

    DIVISION:

    CRM Cardiac Rhythm Management

            

    LOCATION:

    United States > Sunnyvale : 645-647 Almanor Ave

         

    ADDITIONAL LOCATIONS:

    United States > Sylmar : 15900 Valley View Court

         

    WORK SHIFT:

    Standard

         

    TRAVEL:

    Yes, 10 % of the Time

         

    MEDICAL SURVEILLANCE:

    No

         

    SIGNIFICANT WORK ACTIVITIES:

    Continuous sitting for prolonged periods (more than 2 consecutive hours in an 8 hour day)

         

    Abbott is an Equal Opportunity Employer of Minorities/Women/Individuals with Disabilities/Protected Veterans.

         

    EEO is the Law link - English: http://webstorage.abbott.com/common/External/EEO_English.pdf

         

    EEO is the Law link - Espanol: http://webstorage.abbott.com/common/External/EEO_Spanish.pdf

    Numbers & Facts

    LocationSunnyvale, California
    IndustryHealthcare Services
    Company Size10,000 employees or more
    Year Founded1910
    Websitehttp://www.abbott.com/

    About Company

    At Abbott, we are enthusiastic, energetic and committed to doing great work every day. Our employees are passionate about helping to translate science into lasting contributions to health care and the health of people worldwide. At the heart of our organization is our "Promise for Life"—a statement that embodies our company's commitment to employees, shareholders, local communities and the people who depend on our company and products to live healthier lives.

    Vital to our promise is the speed in which we act, respond and deliver. As Abbott employees, we are ready to meet change and challenges head-on. As a result, we are a company that adapts quickly, and through our passion for innovation we are able to continually create a pipeline of products that help improve the length and quality of life around the world.

    We are proud of our rich, more than 120-year history. We continue to be driven to advance leading-edge science and technologies, support diversity, focus on exceptional performance and earn the trust of those we serve.

    Similar Jobs