Lighthouse Technology Services is partnering with our client to fill their Senior (SRE) Site Reliability Engineer (New Environment) position! This is a 3 months with possible extension contract opportunity and will be remote in United States.
This role will be a W2 employee of Lighthouse Technology Services. No C2C or subcontracting arrangements will be considered.
What You'll Be Doing:
Lead the migration of online and mobile banking services from legacy systems to a modern Azure-based environment
Establish and execute release management strategies, defining SLAs and service reliability standards for critical API services
Serve as the technical authority for site reliability engineering, driving performance optimization, resilience, and capacity planning initiatives
Provide production support and incident management, ensuring rapid resolution of issues impacting system stability and availability
Design and implement infrastructure architecture using Azure, Terraform, GitLab CI/CD pipelines, and Artifactory
Support data center migration projects, coordinating deployments and ensuring seamless transitions to new environments
Collaborate with technology management, development teams, and business stakeholders to translate requirements into scalable technical solutions
Define and drive service reliability metrics including SLOs, SLAs, and error budgets for mission-critical applications
Mentor and coach junior engineers on SRE principles, best practices, and infrastructure automation techniques
Participate in on-call rotation to provide off-hours support and ensure 24/7 system availability
What You'll Need to Have:
Senior level experience in infrastructure engineering, systems architecture, or site reliability engineering roles
Expert-level proficiency with Azure cloud services, Terraform, GitLab CI/CD, source code management, and Artifactory
Proven ability to establish release management frameworks, define service reliability standards, and implement error budgets
Strong background in production support and incident management within complex, high-availability environments
Demonstrated experience supporting large-scale migration projects and serving as SME for reliability, performance, and capacity planning
Deep technical expertise in server/client architectures, virtualization technologies, and modern deployment strategies
Excellent communication skills with the ability to influence stakeholders and translate complex technical concepts for diverse audiences
Strong analytical and troubleshooting capabilities with experience resolving performance issues in production environments