Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)
About the Opportunity
Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.
Key Responsibilities
Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
Support log ingestion, index management, lifecycle/retention, and search performance tuning.
Participate in an on-call rotation; support occasional weekend/after-hours needs.
Required Qualifications
US citizenship required; dual citizenship not permitted.
8 years of experience in SRE/DevOps/cloud operations with distributed systems.
Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
Strong Linux expertise (SUSE and Ubuntu).
Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.