The Enterprise Infrastructure SRE team is seeking a seasoned Sr. Site Reliability Engineer to own the reliability, automation, and observability of our data center and network infrastructure. This role bridges network operations and data center facility infrastructure-including bare-metal server reliability, power/cooling systems, and network fabric-with a relentless focus on uptime, scalability, and user experience.
You will minimize manual toil through advanced automation, drive operational excellence through blameless postmortems, and proactively mitigate infrastructure risks. As a senior engineer, you will lead high-impact projects, mentor peers by example, and foster continuous improvement across the global SRE organization.
Enhance Telemetry Platforms: Scale observability, logging, and alerting systems using Grafana, Splunk, and Prometheus
Build Correlated Dashboards: Design visualization tools that connect server health, network telemetry, and facility power/cooling performance across fragmented data sources
Drive Data Insights: Write complex SQL and SPL queries to analyze infrastructure trends, isolate production bottlenecks, and surface environmental health insights via IPMI interfaces
Eliminate Manual Toil: Develop robust automation scripts and tooling to handle hardware incident triage, alert noise reduction, and log correlation
Manage Source of Truth: Maintain and scale Netbox inventory systems, building automated API pipelines to track physical layout, device lifecycles, and cable topologies
Standardize Operational Playbooks: Create and maintain high-quality runbooks, KB articles, and SOPs to enable bot-assisted incident resolution
Lead Incident Response: Participate in high-severity infrastructure on-call rotations, directing rapid triage, mitigation, and root cause analysis (RCA) for server, network, and facility anomalies
Cross-Functional Partnership: Collaborate with architecture, deployment, hardware engineering, and facility operations teams to ensure new implementations are supportable and monitored from day one
Optimize Performance: Continuous monitoring of network and infrastructure performance to execute sustainability and optimization changes
Bachelor s Degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
8+ years of experience in Site Reliability Engineering (SRE), network operations, or data center infrastructure management across highly distributed, large-scale environments
Strong proficiency in Python, Go, or Shell scripting, combined with configuration management frameworks like Salt or Ansible
Observability Expertise: Deep production experience with Prometheus, Grafana, Alert Manager, and Splunk for enterprise metrics and logging
Strong SQL skills (PostgreSQL, MySQL) and a proven track record of consuming and building RESTful APIs to integrate infrastructure tooling
Advanced Linux system fundamentals paired with hands-on experience provisioning, troubleshooting, and managing bare-metal enterprise server architectures
Deep understanding of TCP/UDP, IPv4/IPv6, BGP, EVPN, VxLAN, Segment Routing, and load balancing
Experience managing enterprise hardware vendors like Arista Networks, Juniper Networks, Cisco, Palo Alto Networks firewalls, and F5 load balancers
Practical knowledge of out-of-band management (IPMI), PDU architecture, and data center physical cooling systems (liquid cooling, HVAC, hot/cold aisle containment)
Benefits
Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:
| Location | Austin, TX |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder