Senior Staff DevOps Engineer – Orchestration

Upscale AI
  • US - Headquarters
    30+ days ago

    Job Description

    Why join Upscale AI

    Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.

    We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.

    If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.


    About the role

    Own the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.

    You will build and maintain Kubernetes clusters across on-prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform-driven infrastructure-as-code, and handle secret and certificate rotations.

    You will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end-to-end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.

    Beyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast-moving team.

    What you'll work on

    • Kubernetes cluster lifecycle across hybrid environments: provision, upgrade, scale, and harden clusters running on bare-metal (on-prem customer datacenters) and cloud (AWS/GCP) using kubeadm, Rancher, or equivalent tooling
    • CI/CD pipeline design and ownership: build and maintain pipelines (GitHub Actions or equivalent) that deliver Go microservices, React UI, Helm charts, and edge appliance images from commit to production with automated testing gates
    • Infrastructure-as-code: manage all cloud and on-prem infrastructure through Terraform modules with proper state management, drift detection, and PR-based review workflows
    • Observability stack operations: deploy, tune, and maintain Prometheus (metrics), Grafana (dashboards), Loki (logs), Splunk and Datadog (enterprise monitoring), and Timestream (time-series analytics) — build the dashboards and alerts that give the team real-time visibility into platform and customer-site health
    • Secret and certificate management: automate mTLS certificate rotation across hub-to-edge communication channels, manage Vault or equivalent secret stores, and handle credential lifecycle for multi-tenant deployments
    • Incident response and debugging: own the runbooks, triage production issues across the distributed hub/edge architecture, perform root cause analysis, and drive post-incident reviews that result in real fixes — not just documents
    • Customer site operations: support deployment, upgrade, and troubleshooting of edge appliances running in customer datacenters with varying network constraints and access patterns
    • Internal tooling: build CLI tools, deployment automation, environment provisioners, and operational dashboards that reduce toil and make the engineering team faster
    • Capacity planning and cost optimization: monitor resource utilization across clusters and cloud accounts, right-size workloads, and forecast infrastructure needs as site count grows

    What you bring

    • 8-13 years in SRE, DevOps, or infrastructure engineering roles supporting production distributed systems
    • Deep Kubernetes expertise: cluster administration, networking (CNI, ingress, service mesh), storage (PV/PVC, CSI drivers), RBAC, and troubleshooting pod/node-level issues in both cloud and bare-metal environments
    • Strong Terraform skills with experience managing multi-environment, multi-provider infrastructure at scale
    • Hands-on experience building and operating CI/CD pipelines end-to-end — not just configuring someone else's templates
    • Production experience with at least three of: Prometheus, Grafana, Loki, Splunk, Datadog, Timestream, or comparable observability tools
    • Solid scripting and automation skills in Python, Bash, or Go
    • Working knowledge of Linux systems internals: networking (iptables, DNS, TCP debugging), storage, process management, and performance analysis
    • Experience managing TLS/mTLS certificates, secret rotation, and Vault or equivalent in production
    • Comfort working across cloud (AWS/GCP) and on-prem environments with different constraints and access models
    • Strong debugging instincts — you can follow a problem from a user report through load balancers, ingress, service mesh, application logs, and database queries to root cause

    Nice to have

    • Experience supporting network infrastructure or datacenter automation platforms
    • Helm chart authoring and management for complex multi-service applications
    • Bare-metal Kubernetes provisioning (not just managed EKS/GKE)
    • eBPF-based observability tools (Cilium, Pixie, Hubble)
    • Experience operating Kafka, ClickHouse, ArangoDB, or Redis in production
    • Familiarity with SONiC, network switch management, or ZTP workflows
    • On-call experience with a structured incident management process (PagerDuty, Opsgenie)
    • SOC 2, FedRAMP, or equivalent compliance experience for infrastructure
     
    $222,000 - $250,000 a year

    Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.

    Equal Opportunity

    Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.

    Accessibility & Accommodations

    We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at [email protected]—we’re happy to help. Note: This inbox is only for accommodation requests.

    We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

    Numbers & Facts

    LocationUS - Headquarters
    Websitehttps://upscaleai.com

    Skills

    • Amazon Web Services (AWS)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Bash Scriptingunmatched
    • Capacity Managementunmatched
    • Cloud Computingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • DNS (Domain Name System)unmatched
    • Debugging Skillsunmatched
    • DevOpsunmatched
    • Digital Certificatesunmatched
    • Distributed Computingunmatched
    • Forecastingunmatched
    • Fundingunmatched
    • GCP (Good Clinical Practices)unmatched
    • GitHubunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Linux System Internals/Programmingunmatched
    • Load Balancingunmatched
    • Machine Toolunmatched
    • Metricsunmatched
    • Microservicesunmatched
    • Network Administration/Managementunmatched
    • Network Operations Centerunmatched
    • Network Performance/Analysisunmatched
    • Network Supportunmatched
    • Network Switchingunmatched
    • On Callunmatched
    • Operational Auditunmatched
    • Operational Improvementunmatched
    • Operational Strategyunmatched
    • Operational Supportunmatched
    • Performance Analysisunmatched
    • Problem Solving Skillsunmatched
    • Process Managementunmatched
    • Production Supportunmatched
    • Production Systemsunmatched
    • Public/Media/Press/Analyst Relationsunmatched
    • Python Programming/Scripting Languageunmatched
    • Redisunmatched
    • Reporting Dashboardsunmatched
    • Resource Utilizationunmatched
    • Right-Sizingunmatched
    • Root Cause Analysisunmatched
    • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
    • Scripting (Scripting Languages)unmatched
    • Splunkunmatched
    • TCP (Transmission Control Protocol)unmatched
    • Test Automationunmatched
    • Time Managementunmatched
    • Time Series Analysisunmatched
    • User Interface/Experience (UI/UX)unmatched
    • iptablesunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder