Infrastructure Reliability Operations Administrator- 181480

PeopleSERVE, Inc.

Boston, TX

JOB DETAILS
SKILLS
Agile Programming Methodologies, Artificial Intelligence (AI), Automation, Communication Skills, Computer Operations, Computer Storage Hardware, Dell Computers, DevOps, Disk Management, Hypervisors, IBM AIX Operating System, IT Service Management (ITSM), Identify Issues, Incident Management, Information Technology & Information Systems, Linux Operating System, Maintain Compliance, Metrics, Microsoft Windows Operating System, Network Attached Storage (NAS), On Call, Operating Systems, Operational Improvement, Operational Support, Oracle, Performance Metrics, Problem Solving Skills, Process Development, Process Improvement, Python Programming/Scripting Language, Red Hat Linux Operating System, Reliability Engineering, Root Cause Analysis, Sales Qualification, Security Compliance, Server Hardware, Server Support, Software Patches, Storage Area Network (SAN), Unix Operating Systems, Unix Shell Programming, VMWare, Virtual System Managers, Virtualization
LOCATION
Boston, TX
POSTED
1 day ago
We are seeking a talented Infrastructure Reliability Operations engineer to join our Compute Operations team. In this role, you will collaborate with senior technology professionals, peers across operations teams, and represent Compute Operations during incident calls. You will also work closely with teams based in our India offices on a daily basis. As a key member of the operations team, you will continuously evaluate opportunities for automation, promote DevOps and EngOps practices, and ensure the stability and security of our environment.

The Role

  • Manage and coordinate Linux, Unix, and Windows operating systems
  • Oversee hypervisors and hardware infrastructure
  • Ensure security compliance and execute patching activities
  • Maintain environment stability and support critical server operations
  • Participate in incident management and change execution
  • Contribute to operational KPIs , metrics & observability
  • Drive automation initiatives and promote DevOps/EngOps work
  • Collaborate with global teams and vendors to resolve issues and implement solutions
  • Identify process improvements to improve operational stability
  • Communicate effectively with engineering, operations leaders, and partners

This is a shift-based position:

  • Schedule: Wed–Sat or Sun–Wed (10 hours per day)
  • Flexibility: Occasional coverage on other days may be required

Required Expertise & Skills

  • 5+ years of IT experience across a broad range of technologies, with a focus on server and storage infrastructure
  • Strong knowledge of Linux (Redhat Linux 7, 8, and 9)
  • Disk storage management expertise
  • Experience with virtualization technologies (preferably OLVM)
  • On-call coverage and incident management experience
  • Troubleshooting skills for OS, hardware, and storage issues
  • Shell scripting and Python proficiency; coding experience using AI is a plus
  • Experience working with enterprise-level customers and application teams
  • Windows, VMware, and OLVM (Oracle Linux Virtualization Manager) experience is highly desirable
  • Knowledge of storage subsystems and SAN/NAS/CAS infrastructure is a significant plus
  • Experience crafting and maintaining logging, monitoring, and alerting capabilities using Observability tools

Additional Skills

  • Proven experience in server operations and infrastructure teams
  • Expertise in managing high-severity incident calls
  • Strong understanding of Agile methodology and IT Service Management
  • Comprehensive knowledge of infrastructure tech stack: Linux, AIX, Windows, VMware, OpenStack, Client, Dell, IBM AIX server hardware
  • Ability to identify gaps and drive process improvements
  • Collaboration with vendors (Redhat, Client, Dell, IBM AIX) for root cause analysis and solution implementation
  • Passion for automation, self-service, and self-healing infrastructure
  • Excellent communication and relationship-building skills

About the Company

P

PeopleSERVE, Inc.