RESPONSIBILITIESDeploy, upgrade, operate, maintain, and scale our suite of clusters and servicesCollaborate with engineers to develop automated, full turnkey solutions for silicon simulation workflows to speed up project timelinesManage our underlying infrastructure as code and use modern observability tools to provide a complete picture of cluster and infrastructure healthOperate the continuous integration pipeline, build and release systems, and version control across the environmentIdentify and eliminate performance bottlenecks using measurement and creative engineeringBASIC QUALIFICATIONSBachelor's degree in computer science, information systems, or an engineering discipline; OR 2+ years of professional experience in system administration, high performance computing, or site reliability engineering1+ years of development experience with Bash, Python, and/or other programming languages1+ years of experience with Linux operating systemsPREFERRED SKILLS AND EXPERIENCEFamiliarity with containerization technologies (i.e. Docker, Kubernetes)Knowledge in computer system concepts (computer architecture, computer organization, operating systems and concurrency)Experience with databases and data modeling (e.g.,MySQL,PostgreSQL, SQLite)Networking knowledge of TCP/IPExperience with high performance computing and workload managers (e.g., Slurm, LSF)Experience with Terraform, Ansible, Puppet, or similar automation frameworksExperience building monitoring and alerting as code (e.g., Grafana, Prometheus, custom exporters)Experience with CI/CD automation at scale (e.g., Jenkins, Bamboo, build systems)Experience with infrastructure as code (IaC) tools for managing fleets of serversExperience with using & building REST API clients/serversExperience with enterprise/networked storage automation (e.g., NetApp ONTAP REST API/CLI, NFS)Experience with ASIC design flows and tools (e.g., Cadence, Synopsys, Ansys, Keysight, Siemens)Strong desire to find performance bottlenecks and performance improvement techniquesExcellent communication skills with the ability to communicate with customers, peers, management, etc. in both formal and informal situationsAbility to quickly learn new tools and frameworksInterest in or experience with AI/LLM-assisted tooling (e.g., Grok, Claude Code)ADDITIONAL REQUIREMENTSAbility to work extended hours and weekends as needed to meet critical milestonesCOMPENSATION AND BENEFITSPay Range: Level 1: $125,000.00 - $150,000.00Pay As a Site Reliability Engineer on the Silicon Engineering team you will get the opportunity to design, operate, scale, and automate the high performance computing infrastructure we use to develop the chips powering the world's largest satellite constellation and a global internet service.