AI Experience: Some experience with AI is required
Role Overview
We are looking for a hands-on Software Engineer / Site Reliability Engineer to join a Cloud Platform team. This role combines software engineering with SRE practices, focusing on building and operating scalable, resilient services on AWS.
Key Responsibilities
Design, build, and maintain cloud-native services and platform components on AWS
Apply DevOps/SRE practices including monitoring, alerting, incident response, capacity planning, and on-call support
Develop automation, tooling, and services primarily using Python and/or Go
Manage, tune, and troubleshoot database technologies
Operate and optimize OpenSearch clusters
Automate infrastructure provisioning using Ansible
Improve platform scalability, performance, reliability, and cost efficiency
Participate in postmortems and root-cause analysis
Support AliCloud workloads for multi-cloud initiatives
Required Skills & Experience
AWS: 2–5 years
Database Technologies: 2–5 years
DevOps / SRE: 2–5 years
Python: 2–5 years
OpenSearch: 1+ year
Hands-on experience with production systems, cloud infrastructure, monitoring, troubleshooting, and automation
Some practical experience with AI/AI-related workloads
Nice-to-Have Skills
AliCloud: 1+ year
Ansible: 1+ year
Go / Golang: 1+ year
Multi-cloud experience
Have a Great Day! Warm Regards, Sania Wadhwa Technical Recruiter Contact Details: 609-755-5388 Email: