Looking for a candidate with strong Oracle PL/SQL, Unix batch operations, RCA, and Manhattan UI configuration support.
Role Summary
We are seeking a hands‑on analyst/engineer who is strong in Oracle PL/SQL and Unix/Linux to own daily batch job monitoring, perform root-cause analysis (RCA) for failures and performance issues, and ensure end‑to‑end stability of data processing in a Manhattan (DFIO/SCM) ecosystem. The role includes publishing daily open items, prioritizing work by business impact, and collaborating with DBAs and the Manhattan product team to resolve outliers and configuration issues.
Key Responsibilities
Batch Operations & Monitoring
Monitor and manage daily/overnight batch jobs (Unix shell + SQL/PLSQL steps); validate pre-/post-conditions and recover/rewind missed or skipped steps.
Triage job failures, capture diagnostics, and execute reruns safely with appropriate approvals and documentation.
Root‑Cause Analysis (RCA)
Investigate failures by reviewing Unix logs, job runbooks, and database logs; pinpoint breaking step(s) and document the actual cause (e.g., data, logic, config, capacity).
For each failure/performance incident, produce a clear RCA and corrective & preventive actions (CAPA); drive closure.
Database Analysis (PL/SQL & Performance)
Analyze PL/SQL packages, procedures, and functions implicated in failures; identify defects, data issues, or edge cases.
Diagnose long‑running queries; obtain and interpret execution plans; coordinate with DBA for index, stats, and plan‑stability actions.
Maintain and leverage an up‑to‑date entity‑relationship (ER) understanding of core tables and data flows across DFIO/SCM modules.
Manhattan UI & Configuration Support
For Manhattan UI configuration change requests, validate expected behavior in lower environments; create/execute test scenarios and capture evidence.
For unclear or outlier behaviors, collaborate with the Manhattan product team to confirm product‑level settings and recommended fixes.
Daily Operations Governance
Publish daily open items and status across the team queue; prioritize by business criticality and SLAs; call out blockers and risks.
Adhere to change/incident processes (e.g., ServiceNow: incidents, problems, change requests), maintain accurate runbooks and knowledge base articles.
Quality, Testing & Readiness
Create and execute system test plans for fixes and configuration changes; support release readiness and post‑release validation.
Stakeholder Management
Partner with business users to understand impacts and timelines; provide clear, concise communication on incident status, ETAs, and next steps.