Sr TechOps and SRE Lead (AWS Cloud)\- REMOTE

Simple Solutions

  • Jacksonville, Florida
  • 30 days ago
  • Remote
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • AWS Lambdaunmatched
    • Access Controlunmatched
    • Amazon CloudFrontunmatched
    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Amazon Relational Database Service (RDS)unmatched
    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Architectural Designunmatched
    • Architectural Servicesunmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Benchmarkingunmatched
    • Best Practicesunmatched
    • Budgetingunmatched
    • Cloud Computingunmatched
    • Computer Scienceunmatched
    • Computer Securityunmatched
    • Cost Controlunmatched
    • Crisis Managementunmatched
    • Cross-Functionalunmatched
    • Data Recoveryunmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • High Availabilityunmatched
    • ISO (International Organization for Standardization)unmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Leadershipunmatched
    • Linux Operating Systemunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Network Securityunmatched
    • On Callunmatched
    • Performance Metricsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reliability Engineeringunmatched
    • Resource Utilizationunmatched
    • Root Cause Analysisunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • Software Patchesunmatched
    • Strategic Planningunmatched
    • Systems Reliabilityunmatched
    • Team Lead/Managerunmatched
    • Technical Leadershipunmatched
    • Telephone Skillsunmatched
    • Time Managementunmatched

    Description

    Sr TechOps and SRE Lead (AWS Cloud) - Remote

    Department: Technology / Engineering

    Role Overview

    We are seeking a highly experienced Sr TechOps and SRE Lead  with deep expertise in Cloud to lead our cloud infrastructure, DevOps practices, Site Reliability "Best Practices", and overall operational excellence initiatives. This role is both strategic and hands-on — responsible for designing scalable architectures, improving automation, ensuring system reliability, and leading the TechOps team.

    Key Responsibilities

    1. Architect and manage secure, scalable, and highly available infrastructure on AWS.
    2. Design multi-account AWS environments using AWS Organizations.
    3. Implement VPC architecture, IAM policies, networking, and security best practices.
    4. Oversee EC2, ECS/EKS, Lambda, RDS, S3, CloudFront, and related AWS services.
    5. Optimize AWS cost management and resource utilization.

    Reliability & Production Operations

    1. Implement Site Reliability Engineering (SRE) best practices.
    2. Define SLIs, SLOs, and error budgets.
    3. Manage monitoring and alerting (CloudWatch, Datadog, Prometheus, Grafana).
    4. Lead incident response, root cause analysis (RCA), and postmortems.
    5. Ensure 24/7 uptime and operational resilience.

    Security & Compliance

    1. Implement IAM best practices and least-privilege access controls.
    2. Manage secrets and key management (AWS KMS, Secrets Manager).
    3. Conduct vulnerability management and patching.
    4. Support compliance initiatives (SOC 2, ISO 27001, GDPR as applicable).
    5. Lead disaster recovery planning and backup strategies.

    Leadership & Strategy

    1. Lead and mentor a team of DevOps/TechOps engineers.
    2. Establish operational KPIs and performance benchmarks.
    3. Manage on-call rotations and escalation processes.
    4. Collaborate with Engineering, Product, Security, and Data teams.
    5. Contribute to long-term infrastructure strategy and cloud roadmap.
    <>Required Qualifications
    1. Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
    2. 10+ years in DevOps, Cloud Engineering, or Infrastructure roles.
    3. 5+ years leading technical teams.
    4. Strong hands-on experience with AWS services (EC2, EKS, RDS, S3, IAM, VPC, Lambda).
    5. Deep knowledge of networking, Linux systems, and distributed systems.
    6. Experience with Infrastructure-as-Code (Terraform or CloudFormation).
    7. Strong scripting skills (Python, Bash, or similar).
    8. Experience with containerization (Docker) and Kubernetes (EKS preferred).
    Key Competencies
    1. Strong architectural thinking
    2. Hands-on technical leadership
    3. Crisis and incident management
    4. Strategic planning and execution
    5. Excellent cross-functional communication
    Success Metrics
    1. 99.9%+ production uptime
    2. Reduced deployment lead time
    3. Reduced incident frequency and MTTR
    4. Improved cost efficiency
    5. High-performing and scalable TechOps function




    Numbers & Facts

    LocationJacksonville, Florida (
    Remote
    )
    Websitehttp:\/\/simplesolutions.us

    Similar Jobs