SRE/Devops Engineer
Locations can be any of Tampa, Jersey City or Dallas.
Below is Detailed JD for reference:.
Design & Delivery Partnership: Participate in design reviews, sprint zero, and delivery planning to champion non functional requirements (NFRs) including resiliency, observability, fault tolerance, Holiday and Special days processing, as well as disaster recovery.
Major Release Management Partnership: Collaborate with Major Release Management to ensure each Risk release meets SRE standards for observability and resiliency (SLIs/SLOs, monitoring, knowledge base articles). Ensure releases are subject to required deployment validations.
Monitoring, Observability & AI Enablement: Define and evolve monitoring, alerting, SLIs, and SLOs, leveraging AI/ML driven analytics for anomaly detection, incident correlation, and early risk identification.
Minimizing Application Recovery Time: Make design recommendations that will quick detection of outage conditions and allow the application to recover without manual interventions and/or create a knowledge based guidance for application support team to follow for improved application recovery times. Participate in major incident response / Root Cause analysis to drive continual systemic recovery time improvements.
Automation & Self Healing: Drive automation and intelligent tooling (including AI assisted remediation) to reduce manual toil and improve consistency and recovery times.
Operational Readiness & Risk Management: Attend and present operational readiness with application support (EAS L2) at project management meeting - raise any operational risks and concerns. Test NFRs in UAT environments to validate effectiveness and completeness of operational capabilities. Validate operational readiness prior to release with stakeholders, partner with Embedded Risk and Security teams, and proactively surface and mitigate technology and operational risks.
Capacity & Performance Optimization: Lead capacity planning and performance analysis to ensure Risk platforms scale reliably under high load.
Metrics & Continuous Improvement: Establish KPIs and operational metrics to demonstrate reliability improvements and operational maturity.
People & Culture: Build a strong SRE culture enhanced by AI driven insights across Risk Application Support and Development through mentorship and best practice coaching; leverage approved AI tools to analyze code and collaborate on knowledge base articles, and to accelerate improvements in observability, performance, security, and maintainability.
Qualifications:
Minimum of 8 years of related technical and management experience
Bachelor's degree preferred or equivalent experience
Cloud certifications is a plus
Talents Needed for Success:
Proven experience with SRE or DevOps practices, including CI/CD pipelines, infrastructure as code, and automation frameworks
Strong understanding of monitoring and observability platforms (e.g., Grafana) and experience designing and fine tuning robust monitoring systems
Programming proficiency in one or more languages such as Python, Java, Go, or similar, for automation and tooling development
Familiarity with cloud platforms, containerized environments, and/or hybrid infrastructure models
Experience in financial services, capital markets, or regulated environments
Demonstrated participation in disaster recovery, performance, and resiliency testing
Knowledge of AI concepts, data platforms, messaging systems, and large scale batch or real time processing systems
Strong collaboration skills across technology and business teams
Hands on experience leading and participating in incident and problem management, including root cause analysis
Numbers & Facts
Location
Dallas, TX
Skills
Acceptance Testingunmatched
Analysis Skillsunmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Best Practicesunmatched
Capacity and Performance Managementunmatched
Capital Marketsunmatched
Cloud Computingunmatched
Coachingunmatched
Continuous Deployment/Deliveryunmatched
Continuous Improvementunmatched
Continuous Integrationunmatched
DevOpsunmatched
Disaster Recoveryunmatched
Embedded Systemsunmatched
Financial Servicesunmatched
Go Programming Language (Golang)unmatched
Improvement Metricsunmatched
Incident Managementunmatched
Incident Responseunmatched
Javaunmatched
Knowledge Baseunmatched
Large-Scale Systemsunmatched
Leadershipunmatched
Machine Toolunmatched
Mentoringunmatched
Messaging Technologyunmatched
Metricsunmatched
Operational Improvementunmatched
Performance Analysisunmatched
Performance Metricsunmatched
Performance Testingunmatched
Performance Tuning/Optimizationunmatched
Project/Program Managementunmatched
Python Programming/Scripting Languageunmatched
Release Management/Engineeringunmatched
Reliability Engineeringunmatched
Riskunmatched
Risk Analysisunmatched
Risk Managementunmatched
Root Cause Analysisunmatched
Software Administrationunmatched
Software Developmentunmatched
Team Playerunmatched
Technical Leadershipunmatched
User Documentationunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.