AI Operations & Infrastructure Engineer

Invictus International Consulting, LLC
  • Fort Meade, MD
  • Full-time
  • Quick Apply
5 days ago

Job Description

Title: AI Operations & Infrastructure EngineerLocation: Fort Meade, MDClearance: TS/SCI with a CI Polygraph Job Details:Manage and maintain AI computing platforms, including GPUs and other specialized hardware Install and configure GPU drivers and software Oversee the AI software stack and toolsImplement and manage containerization technologies like Docker and Kubernetes Configure and optimize networking infrastructure for AI workloads, including InfiniBand and Ethernet Manage storage solutions for AI data, considering performance and capacity requirementsDeploy and manage data processing units (DPUs) to accelerate data center workloads Monitor and manage AI cluster health and resource utilization Implement workload management and scheduling tools like Slurm and Kubernetes Ensure efficient power and cooling for AI infrastructure to maintain optimal operating conditionsConfigure high-performance networking solutions for AI and machine learning workloads Optimize network performance to ensure maximum throughput and minimal latency for AI computations Implement and fine-tune network protocols to enhance data transfer speeds and efficiency Integrate NVIDIA networking products with existing AI infrastructure, including servers, GPUs, and storage systems Deploy networking solutions in data centers to ensure seamless connectivity between AI components Diagnose and resolve networking issues impacting AI workloads to maintain optimal system performance Provide technical support and guidance to teams managing AI infrastructure Collaborate with data scientists, researchers, and IT professionals to understand networking requirements and challengesLead deployment and validation of servers and systems for AI enabled platformsConfigure and manage network topologies, BMC, OOB, TPM, power, and coolingInstall, upgrade, and validate GPU-based servers, BlueField DPUs, cables, and transceiversPerform firmware upgrades, hardware validation, and storage setup Configure and administer physical and logical resources, including M IG partitioning and BlueField platforms Install and configure operating systems, cluster software, drivers, containers (Docker), and NGC CLIManage and orchestrate clusters using NVIDIA Base Command Manager, Slurm, Pyxis, Enroot, and Run: AiPerform stress, benchmarking, and burn-in tests using HPL, NCCL, NVIDIA Nemo, and ClusterKit Verify cabling, firmware/software versions, and network signal quality Troubleshoot and resolve hardware, software, storage, and performance faults Replace faulty components and optimize systems for AMD/Intel platforms Monitor, document, and report on cluster health, resource usage, and job performance Ensure secure, efficient, and scalable operation of NVIDIA AI infrastructure, including user access and workload management Requirements:Qualified candidates must hold an active NVIDIA Professional Certification in either AI Networking, AI Infrastructure, or AI OperationsPrior direct, hands-on professional experience administering NVIDIA GPU and data processing unit (DPU) technologies, AI software stacks, and data center environments for high-performance AI workloadsComprehensive expertise in deploying and maintaining AI compute platforms, requiring proficiency in containerization and workload orchestration using Docker, Kubernetes, Slurm, NVIDIA Base Command Manager, and Run:AiMust be capable of configuring physical and logical resources, including Multi-Instance GPU (MIG) partitioning and BlueField platforms, while overseeing critical facility elements such as power, cooling, and storage solutionsThe ability to demonstrate advanced skills in AI networking, specifically configuring and optimizing high-performance InfiniBand and Ethernet fabrics to ensure maximum throughput and minimal latencyCurrent active TS/SCI clearance with a CI PolygraphEqual Opportunity Employer/Veterans/Disabled

Job Posted by ApplicantPro

Numbers & Facts

LocationFort Meade, MD
Job TypeFull-time

Skills

  • Artificial Intelligence (AI)unmatched
  • Benchmarkingunmatched
  • Capacity and Performance Managementunmatched
  • Computer Firmwareunmatched
  • Computer Serversunmatched
  • Computer Storage Hardwareunmatched
  • Computer Systemsunmatched
  • Data Managementunmatched
  • Data Processingunmatched
  • Data Scienceunmatched
  • Device Driversunmatched
  • Dockerunmatched
  • Ethernetunmatched
  • Facilities Managementunmatched
  • GPU (Graphics Processing Unit)unmatched
  • Hardware Installationunmatched
  • Hardware Upgradesunmatched
  • Intel Product Familyunmatched
  • Machine Learningunmatched
  • Network Administration/Managementunmatched
  • Network Configuration Managementunmatched
  • Network Integrationunmatched
  • Network Operations Centerunmatched
  • Network Performance/Analysisunmatched
  • Network Protocolsunmatched
  • Network Topologyunmatched
  • Operations Managementunmatched
  • Performance Managementunmatched
  • Performance Tuning/Optimizationunmatched
  • Problem Solving Skillsunmatched
  • Pyxisunmatched
  • Resource Utilizationunmatched
  • Scientific Researchunmatched
  • Sensitive Compartmented Information (SCI)unmatched
  • System Validationunmatched
  • Systems Administration/Managementunmatched
  • Systems Maintenanceunmatched
  • Team Lead/Managerunmatched
  • Technical Leadershipunmatched
  • Technical Supportunmatched
  • Top Secret Clearanceunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder