Our client, an IT Services and Consulting company, is looking for a Production Support Engineer/Application Support Analyst for their Jersey City, NJ location.
Responsibilities:
Production Support professional to support critical applications and operational dashboards. This role focuses on day-to-day application health monitoring, dashboard checks, high-level incident triage and resolution, and coordination across multiple technology and business partners. The ideal candidate is technically hands-on, comfortable troubleshooting in real time, and experienced in supporting production environments.
Perform daily production support activities, including dashboard monitoring, application checks, and service health verification.
Monitor system performance and availability; identify anomalies and initiate investigation and remediation.
Provide high-level issue triage and resolution, including incident ownership, impact assessment, and escalation when needed.
Troubleshoot and debug application issues using logs, SQL queries, and root cause investigation techniques.
Support and monitor batch processing, including scheduling, failures, reruns, reconciliation, and dependency checks.
Collaborate with multiple teams such as developers, infrastructure, database teams, upstream/downstream application teams, and business stakeholders to drive issue resolution and communication.
Create and maintain runbooks, SOPs, and knowledge articles; contribute to continuous improvement of monitoring and support processes.
Participate in post-incident reviews to document root causes, corrective actions, and preventive measures.
Requirements:
Experience in production support, application support, or operational support for enterprise systems.
Technical skills
Working knowledge of:
Java
Application-level understanding
Ability to debug and interpret stack traces and logs
Pl/sql / sql
Querying databases
Troubleshooting data issues
Validating results
Soap services
Understanding request/response patterns
Knowledge of common failure scenarios
Batch processes
Monitoring and diagnosing failures
Understanding job dependencies
Debugging and log analysis
Production environment troubleshooting
Root cause identification
Incident management
Strong incident management skills.
Prioritization and ownership mindset.
Structured troubleshooting approach.
Clear stakeholder communication.
Preferred qualifications
Experience supporting systems with strict uptime and availability requirements.
Familiarity with monitoring and alerting practices, including dashboards, alerts, health checks, and operational metrics.
Experience with root cause analysis (RCA) methodologies and post-incident documentation.
Understanding of change management and release support within production environments.
Ability to work effectively with multiple teams and partners in a fast-paced environment.
Soft skills
Strong written and verbal communication skills, including timely status updates during incidents.
Detail-oriented with a focus on operational excellence and risk reduction.
Comfortable working under pressure and handling multiple concurrent issues.