HPC AI Systems Administrator

Murray Resources
  • Houston, TX
  • Instant Apply
7 days ago

Job Description

A fast-growing manufacturing company is seeking an HPC AI Systems Administrator to serve as the foundational architect for its growing AI infrastructure. This role will design and maintain a robust compute platform that enables the Development Team to fine-tune and deploy production-level machine learning models, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies. The HPC AI Systems Administrator is responsible for building a secure, scalable, and highly optimized environment to support the company's corporate data initiatives.

Salary + Additional Benefits:
  • $100,000–$140,000 (Dependent on Experience)
  • Medical, Dental, Vision Insurance
  •  401K 

Location: Houston, TX
Type of Position: Direct Hire

Responsibilities:
  • Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.

Requirements:
  • 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure
  • Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context
  • Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management
  • Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems
  • Bachelor's degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience
  • Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions (preferred)
  • Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines (preferred)

Due to the high volume of applications we typically receive, we regret that we are not able to personally respond to all applications. However, if you are invited to take the next step in the process, you will typically be contacted within one week of submitting your application. #LI-DNI
 

Numbers & Facts

LocationHouston, TX

Skills

  • Artificial Intelligence (AI)unmatched
  • CUDA (Compute Unified Device Architecture)unmatched
  • Cloud Architectureunmatched
  • Computer Engineeringunmatched
  • Computer Scienceunmatched
  • Computer Systemsunmatched
  • Dental Insuranceunmatched
  • Device Driversunmatched
  • Distributed Computingunmatched
  • Dockerunmatched
  • Enterprise Protectionunmatched
  • Establish Prioritiesunmatched
  • GPU (Graphics Processing Unit)unmatched
  • High Throughputunmatched
  • Hybrid Cloudunmatched
  • Linux Administrationunmatched
  • Linux Operating Systemunmatched
  • Machine Learningunmatched
  • Manufacturingunmatched
  • Memory Hardwareunmatched
  • Modeling Languagesunmatched
  • Network Architecture/Engineeringunmatched
  • Network Operations Centerunmatched
  • Performance Analysisunmatched
  • Performance Tuning/Optimizationunmatched
  • Production Machiningunmatched
  • Reporting Dashboardsunmatched
  • Resource Managementunmatched
  • Return on Capital Employed (ROCE)unmatched
  • System Integration (SI)unmatched
  • Systems Administration/Managementunmatched
  • Systems Engineeringunmatched
  • Systems Maintenanceunmatched
  • Telemetryunmatched
  • Vendor/Supplier Relationsunmatched
  • Virtualizationunmatched
  • Vision Planunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder