A fast-growing manufacturing company is seeking an HPC AI Systems Administrator to serve as the foundational architect for its growing AI infrastructure. This role will design and maintain a robust compute platform that enables the Development Team to fine-tune and deploy production-level machine learning models, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies. The HPC AI Systems Administrator is responsible for building a secure, scalable, and highly optimized environment to support the company's corporate data initiatives.
Salary + Additional Benefits:
$100,000$140,000 (Dependent on Experience)
Medical, Dental, Vision Insurance
401K
Location: Houston, TX Type of Position: Direct Hire
Responsibilities:
Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Requirements:
3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure
Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context
Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management
Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems
Bachelor's degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience
Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions (preferred)
Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines (preferred)
Due to the high volume of applications we typically receive, we regret that we are not able to personally respond to all applications. However, if you are invited to take the next step in the process, you will typically be contacted within one week of submitting your application. #LI-DNI
Numbers & Facts
Location
Houston, TX
Skills
Artificial Intelligence (AI)unmatched
CUDA (Compute Unified Device Architecture)unmatched
Cloud Architectureunmatched
Computer Engineeringunmatched
Computer Scienceunmatched
Computer Systemsunmatched
Dental Insuranceunmatched
Device Driversunmatched
Distributed Computingunmatched
Dockerunmatched
Enterprise Protectionunmatched
Establish Prioritiesunmatched
GPU (Graphics Processing Unit)unmatched
High Throughputunmatched
Hybrid Cloudunmatched
Linux Administrationunmatched
Linux Operating Systemunmatched
Machine Learningunmatched
Manufacturingunmatched
Memory Hardwareunmatched
Modeling Languagesunmatched
Network Architecture/Engineeringunmatched
Network Operations Centerunmatched
Performance Analysisunmatched
Performance Tuning/Optimizationunmatched
Production Machiningunmatched
Reporting Dashboardsunmatched
Resource Managementunmatched
Return on Capital Employed (ROCE)unmatched
System Integration (SI)unmatched
Systems Administration/Managementunmatched
Systems Engineeringunmatched
Systems Maintenanceunmatched
Telemetryunmatched
Vendor/Supplier Relationsunmatched
Virtualizationunmatched
Vision Planunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.