Our Client, an IT Services and Consultant company, is looking for a DevOps Cloud Engineer – AWS & Alibaba Cloud for their Austin, TX location.
Responsibilities:
Design, implement, operate, and support highly available and scalable infrastructure across AliCloud and AWS environments.
Apply SRE principles to improve availability, latency, performance, capacity, scalability, and operational efficiency of production services.
Define, track, and continuously improve SLIs, SLOs, and operational reliability metrics for critical services.
Build and maintain automated CI/CD pipelines for application and infrastructure deployments using modern DevOps practices.
Automate infrastructure provisioning, configuration, deployments, health checks, maintenance, and repetitive operational activities using scripting and Infrastructure as Code.
Develop automation using Python, Bash/Shell, or equivalent scripting languages to eliminate manual effort and improve operational consistency.
Administer and troubleshoot Linux-based production systems, including OS, processes, services, storage, networking, permissions, patching, system performance, and security.
Support cloud services covering compute, storage, networking, IAM/security, load balancing, monitoring, logging, backup, and disaster recovery.
Implement and maintain Infrastructure as Code using tools such as Terraform and configuration automation using tools such as Ansible.
Work with containerized and cloud-native platforms such as Docker and Kubernetes, including deployment, scaling, troubleshooting, and operational support.
Implement comprehensive monitoring, logging, alerting, and dashboards using tools such as CloudWatch, Prometheus, Grafana, ELK/Splunk, or comparable platforms.
Participate in production incident response, perform systematic troubleshooting and root-cause analysis, and drive corrective and preventive actions through blameless post-incident reviews.
Develop and maintain operational runbooks, automation playbooks, recovery procedures, and technical documentation.
Identify recurring operational issues and eliminate toil through engineering and automation.
Support capacity planning, performance tuning, high availability, failover, backup/recovery, and disaster-recovery readiness.
Collaborate closely with development, infrastructure, security, network, database, and application support teams to improve end-to-end platform reliability.
Promote secure DevOps/SRE practices including least-privilege access, secrets management, vulnerability remediation, patching, and compliance controls.
Participate in change, release, incident, and problem-management processes and support production environments as required.