Want to know if you’re a fit? Upload your resume and let our AI show you.
Skills
Amazon Web Services (AWS)unmatched
Automationunmatched
Benchmarkingunmatched
Capacity Managementunmatched
Chef (Configuration Management)unmatched
Cloud Computingunmatched
Communication Skillsunmatched
Computer Operationsunmatched
Computer Serversunmatched
Continuous Improvementunmatched
Data Partitioningunmatched
Data Recoveryunmatched
DevOpsunmatched
Disaster Recoveryunmatched
Distributed Computingunmatched
Establish Prioritiesunmatched
Forecastingunmatched
Gitunmatched
High Availabilityunmatched
Hosted Searchunmatched
Identify Issuesunmatched
Incident Responseunmatched
Internet Protocolsunmatched
Jenkinsunmatched
Management Strategyunmatched
Memory Hardwareunmatched
Multiplatform/Cross-Platformunmatched
Multitaskingunmatched
On Callunmatched
Operational Auditunmatched
Performance Tuning/Optimizationunmatched
Problem Solving Skillsunmatched
Reliability Engineeringunmatched
Replication and Remote Mirroringunmatched
Retention Programsunmatched
Sales Pipelineunmatched
Software Patchesunmatched
Software as a Service (SaaS)unmatched
SuSE Linuxunmatched
Team Playerunmatched
Test Automationunmatched
Testingunmatched
Ubuntuunmatched
United States Citizenunmatched
Virtualizationunmatched
Web Servicesunmatched
Description
Contract Details
Work Mode: 100% Remote (US-based)
Location: Herndon, VA
Schedule: 40 hours/week
Duration: 08/17/2026 08/16/2027
Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)
About the Opportunity
Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.
Key Responsibilities
Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
Support log ingestion, index management, lifecycle/retention, and search performance tuning.
Participate in an on-call rotation; support occasional weekend/after-hours needs.
Required Qualifications
US citizenship required; dual citizenship not permitted.
8+ years of experience in SRE/DevOps/cloud operations with distributed systems.
Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
Strong Linux expertise (SUSE and Ubuntu).
Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.