HPC Infrastructure and Cluster Engineer

ATR, LLC
  • Springfield, Virginia
  • $180,000–$200,000 Per Year
  • Full-time
2 days ago

Job Description

HPC Infrastructure And Cluster Engineer

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Skills/Qualifications:

Required:

  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
    • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
  • Desired:
    • Familiarity with parallel file systems and high-throughput storage architectures.
    • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
  • Education: Bachelors Degree in Computer Science or a related field
  • ATR is an Equal Opportunity Employer (EOE) who will provide equal employment opportunity to employees and applicants for employment without regard to race, ethnicity, religion, color, sex, pregnancy, national origin, age, veteran status, ancestry, sexual orientation, gender identity or expression, marital status, family structure, genetic information, or mental or physical disability.

Numbers & Facts

LocationSpringfield, Virginia
Job TypeFull-time
Salary$180,000–$200,000 Per Year

Skills

  • Access Controlunmatched
  • Artificial Intelligence (AI)unmatched
  • Bash Scriptingunmatched
  • Broadbandunmatched
  • Computer Scienceunmatched
  • Computer Systemsunmatched
  • File Systemsunmatched
  • GPU (Graphics Processing Unit)unmatched
  • Hardware Administrationunmatched
  • High Availabilityunmatched
  • High Throughputunmatched
  • Identify Issuesunmatched
  • Linux Administrationunmatched
  • Linux Operating Systemunmatched
  • Network Administration/Managementunmatched
  • Network Configuration Managementunmatched
  • Network Operations Centerunmatched
  • Network Supportunmatched
  • Operating Systemsunmatched
  • Performance Engineeringunmatched
  • Performance Tuning/Optimizationunmatched
  • Python Programming/Scripting Languageunmatched
  • Red Hat Linux Operating Systemunmatched
  • Resource Managementunmatched
  • Schedule Developmentunmatched
  • Scripting (Scripting Languages)unmatched
  • Security Complianceunmatched
  • Sensitive Compartmented Information (SCI)unmatched
  • Software Patchesunmatched
  • Storage Architectureunmatched
  • Systems Administration/Managementunmatched
  • Systems Maintenanceunmatched
  • Top Secret Clearanceunmatched
  • Topologyunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder