Site Reliability Engineer

Runloop
  • San Francisco, California
    30+ days ago

    Job Description

    About Runloop

    Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

    The Role

    We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.

    Responsibilities

    • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds

    • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users

    • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind

    • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment

    • Participate in an on-call rotation to support our production systems

    • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing

    • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems

    • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement

    • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience

    • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes

    Qualifications

    • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience

    • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations

    • Strong programming skills in languages like Python or Go

    • Deep expertise in containerization technologies such as Docker and Kubernetes

    • Experience with cloud infrastructure and tools like Terraform and/or Pulumi

    • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog

    • A solid understanding of networking, security, and Linux systems administration

    • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)

    • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity

    • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems

    • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance

    Bonus Points

    • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery

    Benefits

    • Competitive salary and equity

    • Comprehensive health, dental, and vision insurance for employee and dependents

    • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering

    • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks

    Location:

    • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

    Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

    Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.

    Numbers & Facts

    LocationSan Francisco, California

    Skills

    • Amazon Web Services (AWS)unmatched
    • Application Programming Interface (API)unmatched
    • Artificial Intelligence (AI)unmatched
    • Artificial Intelligence (AI) Agentsunmatched
    • Budget Managementunmatched
    • Change Managementunmatched
    • Cloud Computingunmatched
    • Computer Programmingunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Cross-Functionalunmatched
    • Customer Experienceunmatched
    • DevOpsunmatched
    • Distributed Computingunmatched
    • Forecastingunmatched
    • GCP (Good Clinical Practices)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Incident Responseunmatched
    • Leading Edge Technologyunmatched
    • Linux Administrationunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Microsoft Windows Azureunmatched
    • Network Securityunmatched
    • On Callunmatched
    • Problem Solving Skillsunmatched
    • Process Improvementunmatched
    • Product Engineeringunmatched
    • Production Supportunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Release Management/Engineeringunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Software Engineeringunmatched
    • Systems Maintenanceunmatched
    • User Interface Toolsunmatched
    • User Interface/Experience (UI/UX)unmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder