Site Reliability Engineer (OpenSearch) - Remote

Information Consulting Services

  • Herndon, VA
  • 1 day ago
  • Remote
  • Instant Apply
Want to know if you’re a fit?
Upload your resume and let our AI show you.

Skills

  • Amazon Web Services (AWS)unmatched
  • Automationunmatched
  • Benchmarkingunmatched
  • Capacity Managementunmatched
  • Chef (Configuration Management)unmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Computer Operationsunmatched
  • Computer Serversunmatched
  • Continuous Improvementunmatched
  • Data Partitioningunmatched
  • Data Recoveryunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Distributed Computingunmatched
  • Establish Prioritiesunmatched
  • Forecastingunmatched
  • Gitunmatched
  • High Availabilityunmatched
  • Hosted Searchunmatched
  • Identify Issuesunmatched
  • Incident Responseunmatched
  • Internet Protocolsunmatched
  • Jenkinsunmatched
  • Management Strategyunmatched
  • Memory Hardwareunmatched
  • Multiplatform/Cross-Platformunmatched
  • Multitaskingunmatched
  • On Callunmatched
  • Operational Auditunmatched
  • Performance Tuning/Optimizationunmatched
  • Problem Solving Skillsunmatched
  • Reliability Engineeringunmatched
  • Replication and Remote Mirroringunmatched
  • Retention Programsunmatched
  • Sales Pipelineunmatched
  • Software Patchesunmatched
  • Software as a Service (SaaS)unmatched
  • SuSE Linuxunmatched
  • Team Playerunmatched
  • Test Automationunmatched
  • Testingunmatched
  • Ubuntuunmatched
  • United States Citizenunmatched
  • Virtualizationunmatched
  • Web Servicesunmatched

Description

Contract Details

  • Work Mode: 100% Remote (US-based)
  • Location: Herndon, VA
  • Schedule: 40 hours/week
  • Duration: 08/17/2026 08/16/2027
  • Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)

About the Opportunity

Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.

Key Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
  • Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
  • Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
  • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
  • Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
  • Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
  • Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
  • Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
  • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
  • Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
  • Support log ingestion, index management, lifecycle/retention, and search performance tuning.
  • Participate in an on-call rotation; support occasional weekend/after-hours needs.

Required Qualifications

  • US citizenship required; dual citizenship not permitted.
  • 8+ years of experience in SRE/DevOps/cloud operations with distributed systems.
  • Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
  • Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
  • Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
  • Strong Linux expertise (SUSE and Ubuntu).
  • Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
  • Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
  • Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
  • Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.

Preferred Qualifications

  • AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
  • Experience with Cloud Foundry environments.
  • Experience with Jenkins, Chef, and/or Terraform.
  • Experience with Prometheus and Grafana.
  • Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
  • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms.

Work Environment

  • Collaborative, globally distributed team with cross-training opportunities.
  • Participation in an on-call rotation and occasional after-hours/weekend support.

Numbers & Facts

LocationHerndon, VA (
Remote
)

Similar Jobs

See more jobs