Site Reliability Engineer

Alvaria Inc
  • TX
    1 day ago

    Job Description

    About Aspect Software

    Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human-centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.

    We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.

    Position Overview

    We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.

    This person will partner closely with application engineering, platform, security, and product teams to define service-level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem-solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.

    Key Responsibilities

    • Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
    • Partner with engineering teams to define and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and actionable alerts
    • Build automation and self-service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
    • Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
    • Participate in an on-call rotation and lead or support incident response, communication, mitigation, and recovery
    • Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
    • Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
    • Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
    • Strengthen resilience through redundancy, autoscaling, fault-tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
    • Troubleshoot complex issues across applications, infrastructure, networking, databases, and third-party dependencies
    • Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
    • Collaborate across teams and time zones while clearly communicating risk, trade-offs, incident status, and reliability priorities

    Required Qualifications

    • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
    • 4-5+ years of experience in site reliability engineering, DevOps, platform engineering, cloud operations, or production software engineering
    • Hands-on experience operating production workloads in AWS; comparable experience with Azure or GCP is also valuable
    • Proficiency in at least one programming or scripting language such as Python, Go, TypeScript/JavaScript, or C#
    • Experience with infrastructure as code and configuration automation, such as AWS CDK, CloudFormation, Terraform, or similar tools
    • Experience building or supporting CI/CD pipelines and source-control workflows, preferably with GitHub and GitHub Actions
    • Working knowledge of Linux, networking, DNS, load balancing, firewalls, certificates, and common distributed-systems failure modes
    • Experience with observability and monitoring tools such as Datadog, Grafana, CloudWatch or equivalent platforms
    • Understanding of incident management, root-cause analysis, and blameless post-incident practices
    • Strong troubleshooting, documentation, collaboration, and communication skills

    Preferred Qualifications

    • Experience with containerized or serverless architecture, including Kubernetes, Docker, AWS Lambda, or related technologies
    • Experience with AWS services such as API Gateway, CloudFront, S3, IAM, WAF, DynamoDB, Aurora, EventBridge, SQS, Kinesis, Cognito, VPC, Route 53, and Secrets Manager
    • Experience designing multi-region systems, disaster recovery strategies, backup and restoration processes, and business-continuity controls
    • Experience with performance testing, capacity planning, chaos engineering, or reliability testing
    • Familiarity with security, privacy, audit, and compliance requirements for enterprise SaaS products
    • Experience supporting data-intensive, real-time, or high-volume enterprise applications
    • Experience mentoring engineers or leading cross-team reliability initiatives

    Why Join Us?

    • Help shape the reliability practices behind modern workforce technology
    • Work on meaningful, technically challenging systems used by enterprise customers
    • Collaborate with skilled colleagues across engineering, product, security, and operations
    • Influence architecture, tooling, and operational standards as our cloud platforms evolve
    • Access professional development and career-growth opportunities
    • Join a team that values innovation, accountability, inclusion, and measurable customer outcomes

    This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.

    Numbers & Facts

    LocationTX

    Skills

    • AWS Lambdaunmatched
    • Amazon Web Services (AWS)unmatched
    • Architectural Servicesunmatched
    • Artificial Intelligence (AI)unmatched
    • Aspect Workforce Managementunmatched
    • Automationunmatched
    • Autoscalingunmatched
    • Budgetingunmatched
    • Capacity Managementunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Corrective Actionunmatched
    • Customer Experienceunmatched
    • DNS (Domain Name System)unmatched
    • Data Recoveryunmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • Documentationunmatched
    • Enterprise Applicationsunmatched
    • Firewallsunmatched
    • GCP (Good Clinical Practices)unmatched
    • GitHubunmatched
    • Go Programming Language (Golang)unmatched
    • High Availabilityunmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Incident Responseunmatched
    • Infrastructure as a Service (IaaS)unmatched
    • JavaScriptunmatched
    • Linux Operating Systemunmatched
    • Load Balancingunmatched
    • Load Testingunmatched
    • Machine Toolunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Microsoft C# (C Sharp)unmatched
    • Microsoft Windows Azureunmatched
    • Multiplatform/Cross-Platformunmatched
    • On Callunmatched
    • Operational Improvementunmatched
    • Performance Managementunmatched
    • Performance Testingunmatched
    • Privacy Controlsunmatched
    • Problem Solving Skillsunmatched
    • Python Programming/Scripting Languageunmatched
    • Regulatory Complianceunmatched
    • Reliability Engineeringunmatched
    • Reliability Testingunmatched
    • Reporting Dashboardsunmatched
    • Riskunmatched
    • Root Cause Analysisunmatched
    • Scalable System Developmentunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • Software as a Service (SaaS)unmatched
    • Support Documentationunmatched
    • Systems Reliabilityunmatched
    • Team Lead/Managerunmatched
    • Team Playerunmatched
    • Workforce Managementunmatched
    • Workforce Management Softwareunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder