Create frameworks, processes and best practices to be used across Engineering Build meaningful, insightful and actionable SLIs Automate critical portions of engineering processes, to minimize risk and maximize the speed of innovation Manage capacity and performance to help scale our infrastructure both on public and private clouds around the world Deep dive into learning and understanding the mechanism of every application component, and promoting product scalability, stability and performance Continuous delivery, performance fine-tuning and troubleshooting What We're Looking For: The ability to partner and collaborate cross functionally across an engineering organization Strong knowledge of Linux/Unix/BSD internals and experience working with open source software Experience with technologies such as ZooKeeper, with a focus on reliability, automation, operability and performance 5+ years of experience with programming languages (Python, Ruby, etc.) Infrastructure as code a plus (e.g. Automation experience using Python/Ruby/Go – any scripting What You'll Do: Develop software solutions to enable relailbity and operability of large scale distributed systems Build a deep understanding of how systems behave, scale, interact and fail, and use that insight to identity risks and opportunities for remediation Implement monitoring and reporting of our production environments Build tools and automation to eliminate toil and reduce operational overhead.