The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
WHAT THIS CANDIDATE WILL BE DOING
Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
Automate repeatable administration and remediation tasks with Bash and Python.
Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
WHAT WE NEED TO SEE
7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
Strong shell scripting and Python-based automation capability.
Working knowledge of storage and network dependencies affecting Linux host health.
Ability to operate independently in ambiguous, high-severity production situations.
PREFERRED EXPERIENCE
Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.
Experience supporting validation labs or pre-production cluster certification.
Numbers & Facts
Location
Woburn, MA
Skills
Administrative Skillsunmatched
Analysis Skillsunmatched
Ansibleunmatched
Artificial Intelligence (AI)unmatched
Authenticationunmatched
Automationunmatched
Bash Scriptingunmatched
Bootingunmatched
CPU (Central Processing Unit)unmatched
Cloud Computingunmatched
Computer Firmwareunmatched
Computer Maintenanceunmatched
Computer Networksunmatched
Data Administrationunmatched
Data Storageunmatched
File Systemsunmatched
GPU (Graphics Processing Unit)unmatched
Identify Issuesunmatched
Image Managementunmatched
Input/Outputunmatched
Kernel Programmingunmatched
Laboratoryunmatched
Linux Administrationunmatched
Linux Distributionsunmatched
Linux Operating Systemunmatched
Memory Hardwareunmatched
Network Connectivityunmatched
Network Operations Centerunmatched
Operating Systemsunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Red Hat Linux Operating Systemunmatched
Reliability Engineeringunmatched
Telemetryunmatched
Ubuntuunmatched
Unix Shell Programmingunmatched
Vehicle Fleetsunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.