Linux Systems Administrator (HPC experience )
Location: Miami (hybrid onsite 3 days)
Length: 6 mos contract to hire
MUST HAVES ( in order of priority ) Linux Systems Administrator (HPC experience )
AI Clusters and tools, grants
Configuration Management and Provisioning
Storage System experience (Lustre File System)
Backup
Software installation
Windows System Administrator (limited experience okay, as long as strong Linux experience)
Job Description Manages and supports High Performance computing clusters and specialized STEM applications used in scientific computing environments. Provides support in the management, tuning, operating systems, parallel file systems, automating deployment of compute nodes, managing and tuning, login nodes, and HPC scheduling systems. Provides technical support for servers, operating systems and applications, including Cloud Computing, for the entire Enterprise as found in a central IT organization. Designs enterprise environments, installs large scale systems, troubleshoots complex issues involving multiple integrated systems, patch management, performance monitoring and tuning. Provides support for server application environments, virtualization environment, applications and associated infrastructure, such as VDI.
- Administers and supports Linux based enterprise server environments. This includes, but is not limited to, HPC cluster management, VDI, Cloud server management, Linux web servers, NFS servers, VCL Cloud, Linux Satellite servers, VMWare virtualization Environment, and load testing tools. Performs installations and testing patches, applications, performance monitoring and tuning.
- Administers and supports server operating systems and research application environments.
- Troubleshoots and resolves complex system or application issues (includes, but not limited to hardware, storage, operating systems, applications, databases, virtual environments, switches and network issues).
- Ensures proper server and data backups are completed.
- Manages data and systems security/hardening based on Division of IT guidelines.
- Manages HPC job scheduler.
- Manages the back-end support servers or services such as file sharing, database technologies and storage systems as needed.
- Prepares, test, and monitors the distribution of software applications to virtual and physical computer lab environments.
- Assists in the maintenance and continuous improvements of the virtual and physical computer lab environments.
- Performs on-site and remote technical support to faculty, staff, and students.
- Utilizes tools, scripts or programs for automation of routine tasks as required to reduce manual workload, improve response time & ensure system reliability.
- Designs, tests, and recommends new server and application architectures for service improvement. Designs should include performance, scalability, fault tolerance, disaster recovery and fail-over considerations.
- Creates and maintains documentation for all servers, applications, tools, scripts and procedures.
- Assists in the organization and inventory of all hardware and software resources.
- Trains and supervises lower level staff and student employees.
- Provides support for end user service requests as they relate to the applications and systems managed by the Instructional & Research Computing Center (IRCC).
- Performs essential duties during any emergencies, such as hurricanes, storms and/or any other University emergency closing. The employee is expected to be available to report to work as needed during University emergency closings with appropriate notification by department administrator.
Minimum Qualifications
- Bachelor's degree in a technology related field and six (6) years of experience in systems administration, including three (3) years of High Performance computing experience and experience in mentoring or supervising junior systems administrators, OR an equivalent combination of relevant education and/or experience.
Departmental Requirements
- Experience with Linux and Windows server operating systems.
- Understanding of network services.
- Knowledge of hypervisor setup and administration (i.e. Hyper-V, VMware or KVM).
Desired Qualifications
- Certification in Redhat systems administration or other similar Linux distribution.
- Experience and understanding of parallel file systems.
- Experience with programming on Linux systems, especially parallel programming experience.
- Experience running or supporting science, math or engineering applications running on a high performance computing cluster.
- Experience and certification in Cloud Computing.
- Experience with VDI systems.
- Ability to operate Windows server operating systems and technologies.
- Ability to communicate technical information (verbally and written).