7+ years in a Site Reliability Engineering Infrastructure focused role Excellent verbal and written communication skills Automation advocate with a strong sense of ownership - you truly believe in removing operational load via software Strong experience with building and scaling cloud infrastructure and large-scale distributed systems Experience with OpenStack, KVM/hypervisor technologies and Kubernetes Highly proficient in Go or Python Be capable of collaborating and coordinating with multiple distinct engineering teams and mentoring others Experienced with highly distributed unix systemsExperience managing, scaling, and troubleshooting at a planet scale Experience operating large-scale multi-tenant Infrastructure as a Managed service Expert-level proficiency with Infrastructure as Code tools like Chef, Ansible, or Terraform. Design, build and implement innovative solutions around VM orchestration and bare metal provisioning in a highly distributed environment Operate, monitor, and triage all aspects of our production and non-production compute environments Prepare alert handling procedures, runbooks, and collaborate with other SRE teams Participate in on-call rotations to troubleshoot and resolve production issues, minimizing downtime Automate deployment and orchestration of compute infrastructure as well as other routine processes Leverage AI to gain insight across large systems and build tools that can be used in a production environment Interact with and support partner teams, including engineering, QA, and program managementBachelors Degree in Computer Science, an engineering-related field, or equivalent related experience.