Apolis logo

Lead SRE Engineer

Apolis
  • Jersey City, NJ
  • $60–$65 Per Hour
  • Instant Apply
30+ days ago

Job Description

Job Title: Lead SRE Engineer – Strong Java background

Location: Jersey City, NJ – 3 Days onsite role

Long term Project

Job Summary:We are seeking a highly experienced Lead Site Reliability Engineer (SRE) with strong Java development experience to join the technology organization supporting highly available, scalable, resilient, and business-critical applications.

The ideal candidate will have 12+ years of overall technology experience, with strong hands-on expertise in Java, application reliability, production engineering, observability, automation, cloud technologies, CI/CD, incident management, and performance engineering.

The Lead SRE will apply a software engineering mindset to production operations, developing automation and reliability solutions rather than relying solely on traditional infrastructure support. The role will work closely with application developers, architects, DevOps engineers, platform teams, and technology stakeholders to improve system availability, stability, scalability, and operational efficiency.

Banking's technology organization emphasizes application/infrastructure availability, operational stability, system design, Java/Python, resiliency, automation, and strong engineering practices, making this a particularly strong match for a Java-heavy SRE profile.

Key Responsibilities

  • Lead Site Reliability Engineering initiatives for critical enterprise applications and platforms.
  • Develop and maintain Java-based automation, reliability, monitoring, and operational tools.
  • Apply software engineering principles to improve application availability, scalability, resiliency, and performance.
  • Own production stability and participate in incident management, problem management, and root-cause analysis.
  • Design and implement proactive monitoring, alerting, health checks, and automated remediation.
  • Analyze production issues, identify systemic problems, and implement permanent corrective actions.
  • Establish and improve SRE practices, SLOs, SLIs, SLAs, error budgets, and reliability metrics.
  • Build automation to eliminate repetitive manual operational activities.
  • Work closely with Java development teams to improve application reliability and production readiness.
  • Troubleshoot complex Java/JVM, application, API, database, network, and infrastructure-related issues.
  • Perform application performance analysis, including JVM, memory, CPU, thread, garbage collection, latency, and throughput analysis.
  • Design and improve CI/CD pipelines and automated deployment processes.
  • Support highly available applications across cloud and distributed environments.
  • Implement resiliency patterns including fault tolerance, failover, disaster recovery, and capacity planning.
  • Develop dashboards and observability solutions for application and infrastructure health.
  • Participate in production deployments, release management, and post-production validation.
  • Drive automation and continuous improvement across the application lifecycle.
  • Provide technical leadership and mentoring to other engineers.
  • Collaborate with architecture, development, infrastructure, security, and platform engineering teams.

Required Technical Skills

Must Have:

  • 12+ years of IT/software engineering experience
  • Strong hands-on Java development
  • Strong Site Reliability Engineering / Production Engineering experience
  • Java/Spring Boot or enterprise Java application experience
  • Strong understanding of JVM internals and Java application performance
  • Production support and troubleshooting of large-scale applications
  • Linux/Unix
  • REST APIs / Microservices
  • CI/CD
  • Jenkins / GitHub Actions / GitLab CI or similar
  • Kubernetes / Docker
  • Cloud experience — AWS / Azure / GCP
  • Monitoring and observability tools
  • Splunk / ELK or equivalent logging platforms
  • Prometheus / Grafana or equivalent monitoring tools
  • Strong scripting/automation using Python, Shell, or similar
  • Incident management and Root Cause Analysis (RCA)
  • Application performance and capacity management
  • High availability, resiliency, scalability, and disaster recovery concepts
  • Strong SQL/database troubleshooting skills
  • Strong communication and stakeholder-management skills

Numbers & Facts

LocationJersey City, NJ
IndustryComputer/IT Services
Salary$60–$65 Per Hour
Company Size500 to 999 employees
Websitehttps://www.apolisrises.com/

Benefits

Paid Sick Days, Employee Referral Program, Employee Events, Retirement / Pension Plans

About Company

Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.

We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.

Skills

  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Application Performance Managementunmatched
  • Application Programming Interface (API)unmatched
  • Automationunmatched
  • Banking Servicesunmatched
  • Budgetingunmatched
  • CPU (Central Processing Unit)unmatched
  • Capacity Managementunmatched
  • Capacity and Performance Managementunmatched
  • Cloud Applicationsunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Improvementunmatched
  • Continuous Integrationunmatched
  • Corrective Actionunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Dockerunmatched
  • Engineeringunmatched
  • Enterprise Applicationsunmatched
  • Failoverunmatched
  • GCP (Good Clinical Practices)unmatched
  • GitHubunmatched
  • High Availabilityunmatched
  • High Availability Softwareunmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Information Technology Softwareunmatched
  • Javaunmatched
  • Jenkinsunmatched
  • Linux Operating Systemunmatched
  • Memory Hardwareunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Microservicesunmatched
  • Microsoft Windows Azureunmatched
  • Performance Analysisunmatched
  • Performance Engineeringunmatched
  • Problem Solving Skillsunmatched
  • Production Supportunmatched
  • Python Programming/Scripting Languageunmatched
  • REST (Representational State Transfer)unmatched
  • Release Management/Engineeringunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Root Cause Analysisunmatched
  • SQL Databasesunmatched
  • Scripting (Scripting Languages)unmatched
  • Security Infrastructureunmatched
  • Service Level Agreement (SLA)unmatched
  • Software Administrationunmatched
  • Software Developmentunmatched
  • Software Development Lifecycle (SDLC)unmatched
  • Software Engineeringunmatched
  • Splunkunmatched
  • Technical Leadershipunmatched
  • Technical Supportunmatched
  • Unix Operating Systemsunmatched
  • Unix Shell Programmingunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder