SRE Lead engineer

TechDigital
  • Bellevue, WA
    18 days ago

    Job Description

    JD:
    We are looking for a skilled Site Reliability Engineer (SRE) with solid experience in AWS cloud infrastructure to join our growing engineering team. As an SRE, you will ensure our services are reliable, scalable, and performant, while driving operational excellence and infrastructure automation.
    ________________________________________
    Key Responsibilities:
    - Design, implement, and maintain scalable and secure AWS infrastructure.
    - Build and maintain infrastructure as code using Terraform or CloudFormation.
    - Manage Kubernetes (EKS) clusters and containerized workloads.
    - Develop monitoring and alerting solutions using tools like Prometheus, Grafana, and CloudWatch.
    - Support CICD pipelines using tools such as Jenkins, GitHub Actions, or CodePipeline.
    - Participate in incident response, troubleshooting, and root cause analysis.
    - Automate operational tasks through scripting (Python, Bash, etc.).
    - Ensure high availability, reliability, and performance of production systems.
    - Collaborate with development teams to improve system design and architecture.
    ________________________________________
    Required Skills:
    - 3–6 years of SREDevOps experience in a production environment.
    - Strong hands-on experience with AWS services (EC2, S3, IAM, VPC, RDS, Lambda, etc.).
    - Proficient in Terraform or CloudFormation.
    - Experience with Docker and Kubernetes (EKS preferred).
    - Familiarity with Linux system administration and shell scripting.
    - Strong understanding of monitoring, logging, and alerting frameworks.
    - Good knowledge of networking concepts (DNS, TCP/IP, Load Balancing).
    - Strong troubleshooting and incident management skills.

    Numbers & Facts

    LocationBellevue, WA
    IndustryOther/Not Classified
    Company Size100 to 499 employees

    Skills

    • AWS Lambdaunmatched
    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Cloud Computingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • DNS (Domain Name System)unmatched
    • Dockerunmatched
    • GitHubunmatched
    • High Availabilityunmatched
    • High Reliabilityunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Jenkinsunmatched
    • Linux Administrationunmatched
    • Load Balancingunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • System Architectureunmatched
    • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
    • Unix Shell Programmingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder