Experience with monitoring, alerting, and observability tools.
Knowledge of incident management, problem management, and RCA processes.
Experience with automation and scripting using Shell and/or Python.
Working knowledge of PostgreSQL and MySQL databases.
Experience with Git version control.
Understanding of CI/CD concepts and tools such as Jenkins.
Roles & Responsibilities
Design, implement, and maintain highly available and reliable production systems.
Automate operational tasks and infrastructure management using Shell, Python, Ansible, or Terraform.
Manage and support AWS services including EC2, RDS, S3, IAM, VPC, CloudWatch, and related cloud services.
Perform Linux server administration, troubleshooting, patching, and performance tuning.
Monitor application and infrastructure health using tools such as Grafana, Prometheus, CloudWatch, Datadog, Splunk.
Participate in incident management, root cause analysis (RCA), and problem management activities.
Define and maintain SLIs, SLOs, and SLAs to ensure service reliability.
Support PostgreSQL and MySQL databases for operational and basic administration tasks.
Collaborate with development, QA, cloud, and support teams to improve system reliability and deployment processes.
Drive automation, observability, capacity planning, security, and operational best practices.
Participate in on-call support and production issue resolution.
Disclaimer: Diverse Lynx LLC is an Equal Opportunity Employer. All applicants and employees are evaluated without discrimination, based solely on their qualifications, ability, competence and performance. This email and its attachments may contain confidential or proprietary information and is intended only for the recipient(s). If you received this message in error, please disregard it and notify the sender. If you no longer wish to receive our communications, you may unsubscribe at any time.
Security Notice: Our official website is www.diverselynx.com We do not operate or authorize any other websites representing Diverse Lynx LLC.