Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflowsDevelop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overheadDevelop meaningful KPIs and dashboards for business and IT service healthEngineer and implement resilience patterns including HA, DR, and automated failoverPartner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverageConduct performance testing, capacity modeling, forecasting, and right-sizingParticipate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvementsIdentify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutionsDesign, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiencyDesign, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainabilityIdentify gaps in observability coverage and drive engineering solutions to close themCollaborate with architecture and application teams to ensure production readinessLeverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectivenessReduce repeat incidents by engineering permanent fixes and driving continuous improvement. Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practicesStrong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environmentsDemonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environmentAbility to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platformsKnowledge of distributed systems, networking, enterprise infrastructure, and cloud platformsFamiliarity with SRE principles including SLOs, error budgets, observability, and toil reductionAbility to analyze and troubleshoot complex technical systemsPreferred QualificationsExperience in mission-critical, highly available, or regulated environmentsExperience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activitiesKnowledge of ITIL processes and/or SRE best practicesExperience with performance testing, capacity planning, resilience testing, or disaster recovery validation.