We are seeking a highly skilled and motivated Lead Site Reliability Engineer (SRE) to join our team. The Lead SRE will play a critical role in designing, implementing, and maintaining the reliability, scalability, and performance of our cloud-based systems hosted in AWS. You will collaborate closely with software engineers, operations teams, and other stakeholders to enhance system reliability and developer productivity through automation, monitoring, and incident response
What you'll do
Lead a team of SREs consisting of FTEs and contractors
Define and assign the tasks to SREs, review the PRs and provide the feedback
Design and implement scalable, reliable, and secure cloud infrastructure in AWS.
Develop and maintain monitoring, alerting, and dashboarding solutions to ensure system health and uptime.
Automate infrastructure provisioning and configuration management using tools like Terraform and Terragrunt
Implement CI/CD pipelines to streamline deployments and improve development workflows.
Respond to incidents, perform root cause analysis, and implement permanent fixes to prevent recurring issues.
Optimize system performance, reliability, and cost-effectiveness in collaboration with engineering teams.
Drive infrastructure improvements and advocate for best practices in system design and operations.
Establish and manage disaster recovery plans, ensuring system availability during unexpected events.
Qualifications
Minimum Education
BS/BA in Computer Science or equivalent experience