Utah Valley University logo

Site Reliability Engineer III- Operations

Utah Valley University
  • Orem, UT
  • $68,258–$80,304 Per Year
2 days ago

Job Description

Site Reliability Engineer III- Operations

Salary

$68,258.00 - $80,304.00 Annually

Location

Main Campus - Orem

Job Type

FT Exempt Salaried Staff

Job Number

FY2706507

Division

VP Digital Transformation/CIO

Opening Date

08/21/2026

Closing Date

9/4/2026 11:59 PM Mountain

First Review Date

08/28/2026

Required Documents Needed to Apply

Resume

Applicant Support

1-855-524-5627

Support@schooljobs.com

  • Description
  • Benefits
  • Questions

Position Announcement

At Utah Valley University, youll have the opportunity to build and support the technology solutions that help students, faculty, and staff succeed every day. In this role, you will design, implement, and maintain reliable, scalable systems and automation that improve service delivery, strengthen system performance, and support the universitys ongoing digital transformation efforts. Working with modern cloud, infrastructure, and site reliability practices, you will play a key role in ensuring critical services remain secure, available, and responsive for the campus community.

This position offers a dynamic mix of engineering, operations, and innovation. You will collaborate with technical teams and university leaders to enhance system reliability, develop monitoring and automation solutions, support technology upgrades, and resolve complex challenges. Ideal for a technology professional who enjoys continuous learning and problem-solving, this role provides the chance to make a meaningful impact while working in a collaborative environment that values expertise, initiative, and service excellence.

Summary of Responsibilities

  • Performs day-to-day administration, maintenance, upgrades, and operation of existing and recently developed systems, including Virtualization infrastructure on-premises and in the cloud, Microsoft Windows System administration, and Linux operating systems, as well as application systems and technologies.
  • Ensures standard operating procedures, runbooks, disaster recovery, and service catalog definitions are mature and ready for production. Responsible for lifecycle of product set maintenance, availability, reliability, and performance reporting to decision makers and developers. Perform tasks, as needed, to augment work needed on systems by their engineers to ensure timely achievement of project plans and goals. Advocate, contribute, recommend, and facilitate these ever-improving standards and best practices through successful adoption of change within UVU's Digital Transformation department. Ensure that the underlying infrastructure is running smoothly, and that systems and tools are working as expected. Analyze day-to-day functions and the processes of systems and network management software to ensure they are performing within predetermined specifications. In support of core systems availability and reliability, integrate diverse monitoring solutions for emerging and existing IT infrastructure using automation and API tools in on-premises and cloud architectures. Engineer centralized, enterprise-wide alerting and key performance indicators that give timely, actionable information to subject matter experts, stakeholders, and leadership. SRE teams conduct post-incident reviews, documenting findings and acting on lessons learned. Following the incident resolution, the engineer will revisit the issue and determine the cause. Build or optimize the incident lifecycle to bolster the reliability of services. Maintain documentation and runbooks to ensure that teams get information when they need it.
  • Develops operational tools and processes, builds reliable systems, ensures compliance with operational standards, and provides support to operational staff. This can be anything from adjustments to monitoring and alerting to code changes in production. An SRE can be tasked with building a homegrown tool from scratch to help with weaknesses in software delivery or incident response and management. SREs responsibilities include writing and developing code to automate processes, such as analyzing logs, testing production environments, and responding to any issues. Such automation allows developers and engineers to focus their attention on bug fixes and building new features rather than being burdened by the day-to-day operational requirements needed in their projects.
  • Provides leadership, communications, development, engineering, automation, and feedback necessary for enterprise planning and architecture. Timely and responsive work is key for providing what went well or what went badly during a change/incident/problem cycle. Participate in after-hours and weekend on-call rotation and provide training to other on-call staff. Provide remote hands for systems and application administrators that need physical and virtual support within on-premises and cloud facilities.
  • Perform other job-related duties as assigned.

Qualifications / Licenses / Certifications

Graduation from an accredited institution with a bachelors degree in Information Technology or a related field plus three years of work experience in IT; OR a combination of education and experience in a related field totaling seven years.

For Example:

  • Bachelors degree in Information Technology or a related field + 3 years of related work experience.
  • OR two years of completed college coursework in Information Technology or a related field + 5 years of related work experience = 7 years.
  • OR technical certification program or college coursework equivalent to 1 year + 6 years of related work experience = 7 years.
  • OR 7 years of progressively responsible related work experience.

Licenses/Certifications:

  • Site Reliability Engineering (SRE) Professional Certificate
  • Microsoft Certified: Azure Fundamentals, Administrator, Developer, or Associate-level certifications
  • Amazon Web Services (AWS) Cloud Practitioner or Associate-level certifications
  • Information Technology Infrastructure Library (ITIL) certification
  • The Open Group Architecture Framework (TOGAF) certification
  • Docker Certified Associate (DCA)
  • Certified Kubernetes Administrator (CKA)

Knowledge / Skills / Abilities

Knowledge

  • Knowledge of ITIL Change, Incident, and Problem Management.

  • Knowledge of TCP/IP, firewall management, and operating system configuration.

  • Proficient and current knowledge of industry trends, tools, and processes.

  • Knowledge of Agile and iterative development processes (e.g., Scrum and Kanban).

  • Knowledge of automation and containerization technologies such as Docker, Kubernetes, Ansible, Terraform, and SaltStack.

  • Knowledge of ITSM platforms such as Jira Service Management, ServiceNow, or others.

  • Knowledge of Engineering practices: availability, reliability, and scalability, as well as disaster recovery

  • Knowledge of various automation tools, as they are usually responsible for building and integrating software tools to enhance an organizational system's reliability and scalability.

Skills

  • Recognize key design, implementation, and process issues and proactively craft and automate solutions.

  • Skill with system engineering and design for NOC/SOC purposes.

  • Skill with scripting languages such as Perl, PowerShell, Bash, and Python.

  • Skills with most of the common programming languages including JavaScript, HTML5, CSS, JQuery, JSON, and PHP.

  • Skills with the design, implementation, and maintenance of Active Directory and/or LDAP directories.

  • Skills with TCP/IP, application network protocols, firewall management, operating system configuration, anti-virus software, and relational databases.

  • Practical Experience with various Monitoring solutions such as Prometheus, PRTG, Site24x7, TestCafe, Selenium,

  • Splunk, New Relic, Azure Monitor, and AWS CloudWatch.

  • Expertise in the major cloud providers such as Azure, AWS, and Google Cloud.

  • Experience with alert management/on-call tools such as PagerDuty, VictorOps, and Opsgenie.

  • Experience with instant communication and team collaboration platforms like MS Teams, Slack, or Jitsi

  • Proven IT project planning and development skills

Abilities

  • Expert ability to read, write, and interpret technical documentation, runbooks, procedure manuals, and knowledge-base articles pertaining to network systems and application management.

  • Ability to complete Root Cause Analysis (RCA) investigations and write post-incident reports.

  • Ability to improve team practices through code reviews, handoffs of work and incidents.

  • Be on an on-call (PagerDuty) rotation to respond to incidents that impact availability, and provide support for service engineers with customer incidents.

  • Ability to debug production issues and build monitoring that alerts on symptoms rather than on outages.

  • Ability to turn into repeatable actions and into automation.

  • Ability to conduct and direct research into IT issues and products, as required

  • Ability to communicate technical ideas and concepts to a non-technical audience.

EEO Statement:

UVU employment decisions are made on the basis of an applicant's qualifications and ability to perform the job without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, gender expression, age (40 and over), disability, veteran status, pregnancy, childbirth, or pregnancy-related conditions, genetic information, or other bases protected by applicable federal, state, or local law.

Utah Valley University's dedication to exceptional care offers quality service and benefits to employees while staying committed to meeting the needs of a diverse workforce.

UVU is pleased to offer a competitive and comprehensive benefits package that supports employees and their family's overall physical and mental health, protects their income in case of unforeseen illness and life events, and assists in building financial security for retirement and the future.

Highlights from Utah Valley University's benefits package include:

  • Medical network and plan options with low employee premiums
  • Employer HSA contribution for those that elect the university's High-Deductible Health Plan
  • Other tax advantage, reimbursement account options (i.e. Flexible Spending Account & Dependent Care)
  • Dental and vision plan options
  • Incentivized wellness program
  • Employer paid basic life, Accidental Death and Dismemberment (AD&D), and Long Term Disability (LTD) coverage
  • Employee Assistance Program (EAP)
  • 401(a) Defined Contribution Plan with an employer contribution of 14.2% based on employee's compensation (100% vested on first day of full-time employment)
  • Undergraduate tuition remission benefit waiving up to 18 credit hours (each semester) for employees and their eligible dependents (spouses & children dependents unmarried and up to age 26)
  • Generous leave package which includes sick, vacation (staff only), and personal leave; 13 paid holidays; paid medical maternity and parental leave

For more information about the benefits offered, visit https://www.uvu.edu/peopleandculture/benefits.

01

How many years of professional experience do you have administering and supporting enterprise Windows Server environments?

  • Less than 1
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8+

02

How many years of professional experience do you have administering Linux operating systems in a production environment?

  • Less than 1
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8+

03

How many years of experience do you have managing virtualization platforms (e.g., VMware, Hyper-V, Nutanix, Azure VMware Solution)?

  • Less than 1
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8+

04

Please describe your experience supporting cloud infrastructure (e.g., Microsoft Azure, AWS, Google Cloud Platform).

05

How many years of experience do you have performing Site Reliability Engineering (SRE), DevOps, Systems Engineering, or similar responsibilities?

  • Less than 1
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8+

06

Which of the following SRE-related functions have you performed? (Select all that apply.)

  • Monitoring and alerting platform administration
  • Incident response and management
  • Post-incident reviews/root cause analysis
  • Infrastructure automation
  • Performance and reliability engineering
  • Capacity planning
  • Disaster recovery planning and testing
  • Runbook and operational documentation development
  • None of the above

07

Please provide an example of a system reliability, automation, or monitoring solution you designed, implemented, or maintained.

08

Which automation, scripting, or programming languages have you used in a professional environment? (Select all that apply.)

  • PowerShell
  • Python
  • Bash/Shell
  • JavaScript
  • C#
  • Terraform
  • Ansible
  • Other

09

Describe your experience developing scripts, tools, or automation to improve operational efficiency, monitoring, deployment, or incident response processes.

10

Which monitoring and observability platforms have you used? (Select all that apply.)

  • Datadog
  • Splunk
  • Prometheus
  • Grafana
  • Azure Monitor
  • SolarWinds
  • Nagios
  • Zabbix
  • Other

11

Do you have experience participating in an on-call rotation and responding to after-hours production incidents?

  • Yes
  • No

12

Please describe your experience with incident management, troubleshooting critical outages, and performing root cause analysis.

13

Which of the following have you developed or maintained? (Select all that apply.)

  • Standard Operating Procedures (SOPs)
  • Runbooks
  • Disaster Recovery Plans
  • Change Management Documentation
  • Service Catalog Documentation
  • System Architecture Documentation
  • None of the above

14

Please select all licenses and certifications you currently hold. (Select all that apply.)

  • SRE Professional Certificate
  • Microsoft Azure Certification (Fundamentals, Associate, or Expert)
  • AWS Certification (Cloud Practitioner, Associate, Professional, or Specialty)
  • ITIL Certification
  • TOGAF Certification
  • Docker Certified Associate (DCA)
  • Certified Kubernetes Administrator (CKA)
  • Red Hat Certification (RHCSA/RHCE)
  • CompTIA Linux+
  • CompTIA Security+
  • Other Relevant Certification

15

Describe how your education, experience, and technical background have prepared you to support enterprise infrastructure, automation, reliability engineering, and cloud services in a higher education environment.

Required Question

Employer Utah Valley University

Address 800 W. University Parkway

Orem, Utah, 84058

Phone Applicant Support 855-524-5627

Website http://www.uvu.edu

Numbers & Facts

LocationOrem, UT
IndustryEducation
Salary$68,258–$80,304 Per Year
Company Size1,500 to 1,999 employees
Year Founded1941
Websitehttp://www.uvu.edu/

About Company

Utah Valley University was established in 1941 as Central Utah Vocational School (CUVS) with the primary function of providing war production training. CUVS was part of the Provo School District located in south Provo. The institution received a state appropriation in March 1945 of $50,000 to operate for the 1945-1947 biennium. In 1947, the school received funding as a permanent state institution. A new site for the school was acquired on University Avenue in Provo in 1948; in the 1952, the state appropriated funding for the first construction on that site. As enrollments grew, the state acquired over 185 acres in southwest Orem and the first building was completed in 1977. Today, the University’s facilities consist of a combined total of 412 acres with 50 buildings with campuses in Orem, Provo, and Heber City and property in Vineyard and at Thanksgiving Point in Lehi.

Skills

  • Agile Programming Methodologiesunmatched
  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Ansibleunmatched
  • Application Programming Interface (API)unmatched
  • Atlassian JIRAunmatched
  • Automationunmatched
  • Bash Scriptingunmatched
  • Best Practicesunmatched
  • CSS (Cascading Style Sheet)unmatched
  • Capacity Managementunmatched
  • Cloud Architectureunmatched
  • Cloud Computingunmatched
  • Code Reviewsunmatched
  • Communication Skillsunmatched
  • CompTIA Linux+unmatched
  • CompTIA Security+unmatched
  • Debugging Skillsunmatched
  • Dental Insuranceunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Dockerunmatched
  • Document Change Managementunmatched
  • Documentationunmatched
  • Employee Assistance Planunmatched
  • Enterprise Architectureunmatched
  • Firewallsunmatched
  • HTML5unmatched
  • Hardware Virtualizationunmatched
  • Health Planunmatched
  • Higher Educationunmatched
  • IT Service Management (ITSM)unmatched
  • ITIL (IT Infrastructure Library)unmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Industry/Trade Analysisunmatched
  • Information Technology & Information Systemsunmatched
  • Intellectual Property (IP)unmatched
  • JSONunmatched
  • JavaScriptunmatched
  • Kanbanunmatched
  • Knowledge Baseunmatched
  • LDAP (Lightweight Directory Access Protocol)unmatched
  • Leadershipunmatched
  • Linux Administrationunmatched
  • Linux Operating Systemunmatched
  • Maintain Complianceunmatched
  • Manufacturing/Production Testingunmatched
  • Microsoft Active Directoryunmatched
  • Microsoft C# (C Sharp)unmatched
  • Microsoft Certificationsunmatched
  • Microsoft Hyper-Vunmatched
  • Microsoft Windows Azureunmatched
  • Microsoft Windows Serverunmatched
  • Microsoft Windows System Administrationunmatched
  • Nagios Monitoring Toolunmatched
  • Network Management Softwareunmatched
  • Network Operations Centerunmatched
  • Network Systemsunmatched
  • On Callunmatched
  • Operational Improvementunmatched
  • Operational Strategyunmatched
  • Operational Supportunmatched
  • Operations Planningunmatched
  • Operations Processesunmatched
  • PHP Scripting Language (PHP Hypertext Preprocessor)unmatched
  • Performance Analysisunmatched
  • Performance Metricsunmatched
  • Platform as a Service (PaaS)unmatched
  • Problem Solving Skillsunmatched
  • Product Lifecycleunmatched
  • Production Systemsunmatched
  • Programming Languagesunmatched
  • Project Planningunmatched
  • Python Programming/Scripting Languageunmatched
  • Red Hat Linux Operating Systemunmatched
  • Regulatory Complianceunmatched
  • Reliability Engineeringunmatched
  • Research Skillsunmatched
  • Root Cause Analysisunmatched
  • Scripting (Scripting Languages)unmatched
  • Scrum Project Management and Software Developmentunmatched
  • Seleniumunmatched
  • Server Supportunmatched
  • Service Deliveryunmatched
  • ServiceNowunmatched
  • Slackunmatched
  • Software Administrationunmatched
  • Software Engineeringunmatched
  • Splunkunmatched
  • Standard Operating Procedures (SOP)unmatched
  • Systems Administration/Managementunmatched
  • Systems Engineeringunmatched
  • Systems Reliabilityunmatched
  • Systems Scalabilityunmatched
  • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
  • TOGAF - The Open Group Architecture Frameworkunmatched
  • Team Playerunmatched
  • Technical Supportunmatched
  • Technical Writingunmatched
  • Time Managementunmatched
  • Training/Teachingunmatched
  • Tuition Feesunmatched
  • Unix Shell Programmingunmatched
  • VMWareunmatched
  • Virtualizationunmatched
  • Vision Planunmatched
  • jQueryunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder