This hybrid role spans across platform reliability and infrastructure engineering. You'll be instrumental in ensuring high availability, fault tolerance, and performance across internal research and external customers' GPU cluster environments. Responsibilities include automating GPU cluster onboarding, enhancing monitoring, logging, and security systems, and developing new backend features. Required Skills and Certifications:
Proven experience with monitoring tools (e.g., Prometheus, Grafana) and incident management practice.
Strong skills in infrastructure automation with Ansible, Terraform, or similar.
Deep understanding of logging frameworks, alerting systems, and proactive monitoring solutions.
Proficiency in Python for developing automation scripts, REST APIs, and backend support tools.
Hands-on experience with Kubernetes and cloud platforms (GCP preferred).
Knowledge of high-performance networking and real-time systems.
Numbers & Facts
Location
San Francisco, CA
Salary
$100,000–$200,000 Per Year
Skills
Ansibleunmatched
Application Programming Interface (API)unmatched
Automationunmatched
Cloud Computingunmatched
GCP (Good Clinical Practices)unmatched
GPU (Graphics Processing Unit)unmatched
High Availabilityunmatched
Incident Managementunmatched
Multiplatform/Cross-Platformunmatched
Onboardingunmatched
Python Programming/Scripting Languageunmatched
REST (Representational State Transfer)unmatched
Reliability Engineeringunmatched
Scripting (Scripting Languages)unmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.