10+ years of experience in Site Reliability Engineer – OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.
Deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.
Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.
Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
Expertise with Git
Expertise with Concourse, including setup, management, and troubleshooting of new pipelines
Expertise with Linux, specifically SUSE and Ubuntu
Expertise with Kafka, Zookeeper, and Big Data technologies
Expert in development of automation for testing, deployment, scalability, and management of cloud services
Expertise with building, implementing, and/or supporting cloud monitoring tools
Expert knowledge of cloud computing, infrastructure operations, and databases
Expert understanding of web services, networking, virtualization, and internet protocols
Ability to multitask and handle various projects, deadlines, and changing priorities
Excellent communication and prioritization skills
Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
Experience deploying and operating OpenSearch in AWS-based environments
Experience with Cloud Foundry-based environments
Experience with Jenkins, Chef, and/or Terraform
Exposure to and understanding of troubleshooting IP networks and application stacks
Experience with observability tools such as Prometheus and Grafana
Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
Numbers & Facts
Location
Newtown Square, PA (Remote)
Skills
Amazon Web Services (AWS)unmatched
Automationunmatched
Big Dataunmatched
Chef (Configuration Management)unmatched
Cloud Computingunmatched
Communication Skillsunmatched
Computer Operationsunmatched
Data Partitioningunmatched
Data Recoveryunmatched
DevOpsunmatched
Disaster Recoveryunmatched
Distributed Computingunmatched
Establish Prioritiesunmatched
Gitunmatched
High Availabilityunmatched
IP (Internet Protocol)unmatched
Identify Issuesunmatched
Internet Protocolsunmatched
Jenkinsunmatched
Management Strategyunmatched
Multitaskingunmatched
Network Administration/Managementunmatched
Operations Managementunmatched
Performance Tuning/Optimizationunmatched
Production Systemsunmatched
QoS (Quality of Service)unmatched
Query Optimizationunmatched
Reliability Engineeringunmatched
Retention Programsunmatched
Software as a Service (SaaS)unmatched
SuSE Linuxunmatched
Test Automationunmatched
Testingunmatched
Time Managementunmatched
Ubuntuunmatched
Virtualizationunmatched
Web Servicesunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.