Site Reliability Engineer - Remote

DivIHN Integration Inc
  • NULL, NULL
  • Remote
  • $72.85
  • Quick Apply
7 days ago

Job Description

For further inquiries about this opportunity, please contact our Talent Specialist, Marshelin at (224) 507 1280
Title: Site Reliability Engineer - Remote
Location: Corning, NY -
Remote
Duration: 6 Months with possible extension based on business need
Schedule: Candidates can live anywhere in the US but must be able to work 8 AM 5 PM/9 AM 6 PM EST
Only W2 candidates are eligible for this position. Third-party or C2C candidates will not be considered.
Job Description:
Primary Purpose of the Role
The Site Reliability Engineer will support and manage Clients Kubernetes platform infrastructure that supports scientific and engineering applications.
Role Overview
  • Join Client's Model Operations and Deployment Engineering team at their flagship research facility, where your work will directly support groundbreaking materials science innovations. In this role, you will help maintain, enhance, and evolve the Kubernetes platforms that enable scientific and engineering teams to deploy, operate, and scale critical applications across both on-premises and cloud environments.
  • Client is looking for an experienced contract Site Reliability Engineer who can strengthen our team's platform engineering and operational capabilities. You will play a key role in supporting Kubernetes infrastructure managed through Rancher, improving system reliability and automation, and advancing infrastructure-as-code and GitOps practices across our environment.
This role focuses on:
  • Managing and maintaining Kubernetes platforms
  • Ensuring platform reliability, stability, and uptime
  • Supporting both on-premises and cloud-based Kubernetes environments
  • Managing infrastructure that hosts critical business and research applications
  • Supporting platform operations rather than application development
  • Helping the team expand its cloud migration initiatives
  • This is NOT a Software Developer role.
Key Responsibilities
  • Platform Operations: Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.
  • Cluster Management: Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
  • Linux Systems Administration: Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
  • Infrastructure as Code: Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
  • GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
  • Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
  • Continuous Improvement: Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.
Top requirements:
  • Bachelors is preferred, but not required.
  • Minimum of 5 years professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
  • Candidates must have 5 years strong system admininstration with Linux! Rancher for Kubernetes experience is a must.

Top Required Skills (Must-Have)

1. Linux Administration (Highest Priority)

Experience Level

  • Preferred: 10+ years Linux experience
  • Acceptable: Minimum 5+ years strong Linux administration experience

2. Kubernetes Administration

Experience

  • Minimum 3-5 years Kubernetes administration
  • Production environment experience required

3. Rancher

4. Infrastructure as Code

5. GitOps
    Qualifications Required:
    Education prefrred:
    • BS in Computer Science, Software Engineering, Information Technology, or related field preferred; or equivalent professional experience.
    Experience:
    • 5+ years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
    • Hands-on experience operating and supporting Kubernetes platforms in production environments.
    • Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
    • Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
    • Experience with Rancher for Kubernetes cluster management and platform operations.
    • Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management.
    • Demonstrated success working in Agile teams (Scrum, Kanban).
    Technical Skills:
    • Kubernetes: Cluster operations, upgrades, networking, storage, troubleshooting, and workload support.
    • Platform Management: Rancher or similar Kubernetes management platforms.
    • Linux: Advanced administration of Linux/Unix systems.
    • Infrastructure as Code: Strong IaC experience; Cluster API (CAPI) preferred.
    • GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices.
    • Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling.
    Preferred:
    • Experience with hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, GCP).
    • Experience with Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
    • Familiarity with security best practices for Kubernetes and Linux platforms.
    • Experience supporting scientific research environments, high-performance computing, or computational science workflows.
    • Knowledge of CI/CD pipeline development and platform automation patterns.

    Certifications Preferred

    • Certified Kubernetes Administrator (CKA)
    This is the most valued certification for this role.
    Candidates with strong experience can still be selected without certifications.

    Skills That Are Not Priorities

    • Terraform
    Interview Process:
    • Round 1: Teams Interview
    • Round 2: Teams Interview
    • Additional Team Members

    About us:
    DivIHN, the 'IT Asset Performance Services' organization, provides Professional Consulting, Custom Projects, and Professional Resource Augmentation services to clients in the Mid-West and beyond. The strategic characteristics of the organization are Standardization, Specialization, and Collaboration.

    DivIHN is an equal opportunity employer. DivIHN does not and shall not discriminate against any employee or qualified applicant on the basis of race, color, religion (creed), gender, gender expression, age, national origin (ancestry), disability, marital status, sexual orientation, or military status.

    Numbers & Facts

    LocationNULL, NULL (
    Remote
    )
    Salary$72.85

    Skills

    • Administrative Skillsunmatched
    • Agile Programming Methodologiesunmatched
    • Amazon Web Services (AWS)unmatched
    • Application Programming Interface (API)unmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Best Practicesunmatched
    • Business Strategyunmatched
    • Cloud Computingunmatched
    • Computer Scienceunmatched
    • Consultingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Customer Support/Serviceunmatched
    • DevOpsunmatched
    • Ecosystemsunmatched
    • Engineering Softwareunmatched
    • GCP (Good Clinical Practices)unmatched
    • Gitunmatched
    • Identify Issuesunmatched
    • Information Technology & Information Systemsunmatched
    • Kanbanunmatched
    • Linux Administrationunmatched
    • Linux Operating Systemunmatched
    • Machine Toolunmatched
    • Material Scienceunmatched
    • Microsoft Windows Azureunmatched
    • Network Configuration Managementunmatched
    • Network Performance/Analysisunmatched
    • Network Systemsunmatched
    • Operational Supportunmatched
    • Operations Managementunmatched
    • Performance Tuning/Optimizationunmatched
    • Production Systemsunmatched
    • Professional Servicesunmatched
    • Public Cloudunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Sales Managementunmatched
    • Scientific Researchunmatched
    • Scripting (Scripting Languages)unmatched
    • Scrum Project Management and Software Developmentunmatched
    • Service Deliveryunmatched
    • Software Configuration Managementunmatched
    • Software Developmentunmatched
    • Software Engineeringunmatched
    • Source Code/Configuration Management (SCM)unmatched
    • Systems Administration/Managementunmatched
    • Systems Analysisunmatched
    • Systems Engineeringunmatched
    • Systems Reliabilityunmatched
    • Team Playerunmatched
    • Technical Supportunmatched
    • Unix Operating Systemsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder