Job Description
Site Reliability Engineer (SRE)
- Location: Minnetonka Mills, MN
- Duration: 6 months
- Experience Required: 6–8 years
- Relevant Experience:
- 5+ years of experience as an SRE
- Strong knowledge of AWS and Kubernetes
- 3+ years of Telemetry experience
Technical / Functional Skills
- Deep knowledge of platform engineering
- Strong experience with:
- AWS
- Kubernetes
- Cloud platforms
- Act as a bridge between:
- Platform Engineering
- Application Engineering
- Application SRE teams
- Strong expertise in Telemetry
- Proactively identify and eliminate single points of failure
- Establish and maintain application configuration standards
- Work with enterprise teams across:
- Platform
- Network
- Storage
- Infrastructure
- Collaborate with external vendor teams
- Support:
- Upgrade planning
- Platform migrations
- Application configuration standards
- High-availability (HA) initiatives
- Strong monitoring and observability experience with:
Key Responsibilities
- Build and maintain telemetry solutions for business applications
- Support engineering teams with telemetry and observability
- Maintain a close working relationship with Engineering Leaders and Service Directors
- Support a suite of applications and services
- Focus on reliability engineering and system availability
- Identify and eliminate single points of failure
- Monitor infrastructure capacity
- Perform infrastructure tuning and optimization
- Analyze telemetry and trend data
- Identify reliability and performance issues
- Define and measure:
- SLA — Service Level Agreement
- SLO — Service Level Objective
- SLI — Service Level Indicator
Generic / Managerial Skills
- Comfortable interacting with all levels of an organization
- Strong vendor and third-party communication skills
- Architecture and technical experience
- Strong DevOps background
- Ability to work effectively in a matrix environment
- Proven ability to achieve goals collaboratively
- Strong influencing and stakeholder-management skills
- Experience with integration testing
- Experience with test automation practices
- Understanding of LeSS (Large-Scale Scrum) Framework
- Product-focused mindset
Key Skills / Keywords
- Site Reliability Engineering (SRE)
- AWS
- Kubernetes
- Platform Engineering
- Telemetry
- Observability
- High Availability
- Reliability Engineering
- Grafana
- Datadog
- Splunk
- EAPM
- DevOps
- Infrastructure Capacity & Tuning
- SLA / SLO / SLI
- Platform Migration
- Application Configuration
- Integration Testing
- Test Automation
- LeSS Framework
Numbers & Facts
| Location | Minnetonka Mills, MN |
| Salary | $32.57–$43.18 Per Hour |
Skills
Amazon Web Services (AWS)unmatched
Business Solutionsunmatched
Cloud Computingunmatched
Communication Skillsunmatched
DevOpsunmatched
High Availabilityunmatched
Integration Testingunmatched
Quality Assurance Methodologyunmatched
Reliability Engineeringunmatched
Scrum Project Management and Software Developmentunmatched
Service Level Agreement (SLA)unmatched
Software Administrationunmatched
Software Configuration Managementunmatched
Splunkunmatched
Systems Reliabilityunmatched
Team Playerunmatched
Technical Supportunmatched
Telemetryunmatched
Test Automationunmatched
Trend Analysisunmatched
Level up your application
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder