HPC Engineer

Institute of Foundation Models

  • Sunnyvale, California
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Web Services (AWS)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Cloud Computingunmatched
    • Computer Engineeringunmatched
    • Computer Scienceunmatched
    • Computer Systemsunmatched
    • DevOpsunmatched
    • Disability Insuranceunmatched
    • Documentationunmatched
    • Electrical Engineeringunmatched
    • GCP (Good Clinical Practices)unmatched
    • GPU (Graphics Processing Unit)unmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Information Technology & Information Systemsunmatched
    • Life Insuranceunmatched
    • Linux Administrationunmatched
    • Linux Operating Systemunmatched
    • Mathematicsunmatched
    • Metricsunmatched
    • Microsoft Windows Azureunmatched
    • Operational Supportunmatched
    • Performance Analysisunmatched
    • Physicsunmatched
    • Production Supportunmatched
    • Python Programming/Scripting Languageunmatched
    • Research Skillsunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • Supercomputingunmatched

    Description

    About MBZUAI
    The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.

    Position Summary
    This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

    Responsibilities

    • Monitor health, performance, and availability of large-scale GPU clusters.
    • Respond to incidents and perform first-level triage.
    • Support researchers and troubleshoot job failures.
    • Execute operational runbooks and recovery procedures.
    • Validate cluster deployments, upgrades, and maintenance activities.
    • Track infrastructure utilization and operational metrics.
    • Develop automation and monitoring tools.
    • Contribute to documentation and reporting.

    Education

    Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

    Experience

    • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
    • Strong Linux troubleshooting skills.
    • Experience with scripting using Python or Bash.

    Preferred Qualifications

    • Slurm.
    • GPU infrastructure.
    • AWS, Azure, or GCP.
    • Grafana, Prometheus, Datadog, or similar tools.
    • Containers and Kubernetes.
    • AI/ML infrastructure exposure.
    • Research computing environments.
    $150,000 - $300,000 a year
    Salary Range
     
    The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.  
    Benefits Include
    *Comprehensive medical, dental, and vision benefits 
     *Bonus
    *401K Plan
    *Generous paid time off, sick leave and holidays
    *Paid Parental Leave
    *Employee Assistance Program
    *Life insurance and disability
     

    Numbers & Facts

    LocationSunnyvale, California

    Similar Jobs

    See more jobs