Apache Kafka, Artificial Intelligence (AI), Automation, Budgeting, Continuous Deployment/Delivery, Continuous Integration, Metrics, Operational Improvement, Problem Solving Skills, Quality Management, Reliability Engineering, Reporting Dashboards, Spring Framework
Mandatory Skills:
1. Java Spring boot 2. Apache Kafka 3. Dev Ops 4. CI/CD automation
Years of experience required: 8-10
Job Description:
Automation & Efficiency
Automate the top 5 high-volume support and request types
Build self-service and agent-driven solutions to reduce manual work
Harden operational workflows for consistency, auditability, and resilience
Implement auto-retry and backoff for recurring failure patterns
Reliability Engineering
Define and manage Service Level Objectives (SLOs) for critical services and batch processes
Apply error budget concepts to guide reliability and release decisions
Improve batch reliability through standardized recovery patterns and monitoring
Observability & Metrics
Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-Service
Develop and expand runbooks for key production scenarios
Convert runbooks into automated remediation workflows
Enable self-service for repeat operational requests
Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent Operations
Implement self-healing capabilities to minimize manual intervention
Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
Leverage automation and AI to resolve recurring issues with minimal human involvement