Description:
We are seeking an experienced AWS Platform Support Lead to oversee cloud platform operations, production support, incident management, and infrastructure reliability across AWS environments. The ideal candidate will lead a team of support engineers, drive operational excellence, ensure platform stability and security, and act as the primary escalation point for critical incidents. This role requires strong expertise in AWS cloud services, infrastructure operations, automation, monitoring, and ITIL-based service management processes.
Key Responsibilities
Technical Leadership
Platform Operations & Support
Cloud Infrastructure Management
Manage and support AWS services including:
Monitoring & Reliability
Automation & DevOps
o Terraform
o CloudFormation
o Ansible
o Python
o Shell Scripting
Security & Compliance
Stakeholder Management
Responsibilities:
Manage and support AWS infrastructure across multiple environments. Ensure high availability, scalability, performance, and reliability of cloud platforms. Oversee incident, problem, change, and release management activities. Drive major incident management and coordinate recovery efforts. Perform Root Cause Analysis (RCA) and implement preventive actions. Establish proactive monitoring and alerting mechanisms. Define and track platform SLAs, KPIs, and operational metrics. Analyze trends and recommend improvements for platform stability. Ensure capacity planning and performance optimization. Drive automation of operational activities using Infrastructure as Code (IaC). Support CI/CD processes and release deployments. Implement operational automation using Terraform, CloudFormation, Ansible, Python, Shell Scripting. Ensure cloud environments adhere to security policies and compliance requirements. Review IAM configurations, access controls, and overall security posture. Collaborate with security teams to remediate vulnerabilities and audit findings. Partner with application, DevOps, security, network, and business teams. Provide regular operational reports and service review updates. Lead customer and stakeholder communications during major incidents and service reviews. Manage and coach a team of cloud support engineers. Conduct performance reviews and development planning. Drive shift governance and support coverage planning. Ensure adherence to operational SLAs and OLAs. Lead major incident bridges and stakeholder communications. Drive continuous service improvement initiatives.
Qualifications:
Bachelor's degree in Computer Science, Information Technology, or a related discipline. 8%2B years of IT infrastructure and production support experience. 5%2B years of hands-on AWS cloud operations experience. Experience leading platform support teams in a production environment. Strong understanding of cloud networking, security, and infrastructure architecture. Experience working within ITIL-based service management environments. Excellent communication, leadership, and stakeholder management skills. AWS Certified Solutions Architect Professional, AWS Certified SysOps Administrator Associate, AWS Certified DevOps Engineer Professional, AWS Certified Security Specialty, ITIL Foundation Certification.
8%2B years of work experience with Amazon Web Services (AWS)
5%2B years of work experience with Cloud Infrastructure
3%2B years of work experience with Terraform
Shift Details:
global 24x7 operations
Hiring Terms:
No travel required
Clearance: Other
Applicants must be legally authorized to work in the United States.
Additional Details:
Themesoft Inc is an Equal Employment Opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other status protected by applicable federal, state, or local law.