Mandatory Skills- Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management
Key Responsibilities
" Monitor, maintain, and improve the reliability, availability, and performance of production systems.
" Design and implement monitoring, alerting, logging, and observability solutions.
" Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
" Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
" Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
" Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
" Perform capacity planning, performance tuning, and scalability assessments.
" Support CI/CD pipelines and deployment automation initiatives.
" Implement high-availability, disaster recovery, and failover strategies.
Required Skills
" Strong experience with Linux/Unix administration.
" Proficiency in scripting languages such as Python, Shell, or PowerShell.
" Hands-on experience with cloud platforms (AWS, Azure, or GCP).
" Experience with containerization technologies such as Docker and Kubernetes.
" Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
" Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
" Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
" Strong troubleshooting, debugging, and problem-solving skills.
" Understanding of networking, security, and distributed systems concepts.
Experience
" 5 10+ years of overall IT experience.
" 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.