Sr Staff Site Reliability Engineer

Archer

  • San Jose, CA
  • Today
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Aerospace and Defenseunmatched
    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Best Practicesunmatched
    • Cloud Applicationsunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Programmingunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Data Managementunmatched
    • DevOpsunmatched
    • Distributed Computingunmatched
    • Diversityunmatched
    • Dockerunmatched
    • Equal Employment Opportunity (EEO)unmatched
    • HIPAA (Health Insurance Portability and Accountability Act)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Jenkinsunmatched
    • Multiplatform/Cross-Platformunmatched
    • On Callunmatched
    • Operating Systemsunmatched
    • Problem Solving Skillsunmatched
    • Process Improvementunmatched
    • Production Supportunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Assurance Methodologyunmatched
    • Release Management/Engineeringunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Sales Pipelineunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Development Lifecycle (SDLC)unmatched
    • System Architectureunmatched
    • Team Playerunmatched
    • Test Automationunmatched
    • Windows PowerShellunmatched

    Description

    Overview Archer is an aerospace company based in San Jose, California building an all-electric vertical takeoff and landing aircraft with a mission to advance sustainable air mobility. We design, manufacture, and operate an all-electric aircraft capable of carrying four passengers with minimal noise. We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for the reliability, scalability, performance, and security of our core systems and services, designing, implementing, and maintaining robust infrastructure and automation solutions.Responsibilities Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.Implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS).Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems.Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms.Drive the continuous improvement of CI/CD pipelines, promoting best practices for automated testing, deployment, and release management.Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations.Implement and enforce robust security practices across our infrastructure, applications, and data.Design and maintain Docker-based containerization solutions for our applications.Develop and maintain automation scripts and tools using Python, Bash, and PowerShell.Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception.Troubleshoot complex production issues across various layers of the stack, identifying root causes and implementing preventative measures.Participate in on-call rotations to support production systems.Qualifications 12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence.Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting.Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar.Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark).Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD).Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles.Comprehensive understanding of security best practices for cloud environments, applications, and data.Proficiency in Docker for containerization and orchestration.Advanced scripting and programming skills in Python, Bash, and PowerShell.Solid understanding of networking concepts, distributed systems, and operating systems.Excellent problem-solving, analytical, and communication skills.Ability to work independently and as part of a highly collaborative team.Bachelor\'s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.Preferred Qualifications Experience with other Kubernetes distributions or cloud providers.Familiarity with compliance frameworks (e.g., SOC 2, HIPAA, GDPR).Certifications in AWS, Kubernetes, or other relevant technologies.Archer is committed to equal opportunity employment and diversity in the workplace. All aspects of employment are decided on the basis of merit, qualifications, and business needs. We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state, or local laws.#J-18808-Ljbffr

    Numbers & Facts

    LocationSan Jose, CA

    Similar Jobs