A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization. This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes.
The ideal candidate will have a strong background in SRE, DevOps, or Platform Engineering and enjoy balancing hands-on operational support with long-term engineering improvements. You'll work closely with engineering, data, infrastructure, and security teams to enhance platform reliability, automate manual processes, and drive modernization efforts across the environment.
Responsibilities
Provide operational ownership of workflow orchestration and enterprise scheduling platforms, including Apache Airflow and similar technologies.
Act as an escalation point for platform-related incidents, troubleshooting complex production issues and driving root cause analysis through resolution.
Partner with application and data teams to resolve workflow failures, dependency issues, scheduling conflicts, and performance bottlenecks.
Develop and maintain reliability standards, service objectives, monitoring strategies, and operational best practices.
Build automation and tooling that reduce manual effort and improve the overall user experience for engineering teams.
Design and implement observability solutions utilizing metrics, dashboards, alerting, logging, and performance monitoring.
Support platform lifecycle management, including upgrades, patching, configuration management, and security remediation.
Contribute to infrastructure modernization initiatives involving cloud services, containerization, platform migrations, and deployment automation.
Develop and maintain Infrastructure-as-Code solutions using tools such as Terraform, Helm, Ansible, and related technologies.
Participate in architectural discussions, platform roadmap planning, and engineering standards development.
Maintain operational documentation, procedures, and on-call readiness for supported environments.
Required Experience
Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or DevOps-focused roles.
Strong production experience supporting Apache Airflow environments.
Experience with distributed Airflow deployments, including Celery and/or Kubernetes executors.
Experience supporting enterprise workload automation and job scheduling platforms such as Automic/UC4, Control-M, or comparable technologies.
Strong Linux administration skills with working knowledge of Windows-based environments.
Experience supporting cloud infrastructure, preferably within AWS environments.
Proficiency with Python and scripting for automation, tooling, and operational efficiency.
Experience working with Kubernetes, Docker, CI/CD pipelines, and modern deployment methodologies.
Strong understanding of monitoring, logging, tracing, and observability concepts.
Experience with tools such as Grafana, Prometheus, ELK, or comparable monitoring platforms.
Proven ability to manage production incidents and communicate effectively during high-priority situations.
Strong automation mindset with a focus on improving efficiency and reducing operational overhead.