Requisition ID: 102959-1
Title: Platform Reliability Engineer
Duration: Morristown NJ
Salary Range: $37- $40 an hour on W2 /C2C
Job Description:
The Platform Reliability Engineer operates and improves enterprise platforms that support IT operations, automation, observability, AI-enabled workflows, and business-critical services. The role ensures reliable, secure, scalable, and well-governed platforms through strong administration, SRE practices, release discipline, vendor coordination, self-service enablement, and continuous improvement.
Core Responsibilities
Administer platforms, access, integrations, configuration, onboarding, agent registration, inventory, dependency mapping, and lifecycle records.
Maintain platform availability, performance, capacity, recoverability, observability, security posture, and operational readiness.
Plan and execute upgrades, patching, releases, regression testing, rollback readiness, and post-release validation.
Manage vendor escalations, licensing consumption, platform utilization, plugin governance, roadmap alignment, and service improvement actions.
Enable self-service through workflows, catalogs, knowledge articles, automation, runbooks, and simplified support processes.
Support ITSM practices for incidents, changes, problems, requests, assets, configuration, releases, and knowledge management.
Required Skills and Qualifications
3 7 years of experience in platform engineering, IT operations, production support, DevOps, SRE, service management, or related technology roles.
Strong knowledge of platform health, release management, patching, upgrades, regression testing, rollback planning, and production readiness.
Hands-on understanding of SRE concepts such as SLIs, SLOs, error budgets, incident command, postmortems, toil reduction, and reliability reviews.
Experience troubleshooting production issues using logs, metrics, traces, alerts, telemetry, dependency maps, and configuration data.
Knowledge of observability, dashboards, alert tuning, anomaly detection, capacity management, performance analysis, and invocation frequency trends.
Strong documentation, automation mindset, analytical thinking, communication, ownership, and cross-functional collaboration skills.
SRE Focus Areas
Reliability and Resilience: Improve availability, recoverability, performance, capacity health, and operational readiness.
Observability and Incident Response: Build dashboards and alerts, reduce noise, improve detection, support triage, escalation, restoration, RCA, and postmortems.
Automation and Release Reliability: Reduce toil through self-service and workflow automation while supporting release gates, validation, rollback readiness, and post-implementation reviews.
Preferred Technical Experience
Experience with observability, automation, AI/agent platforms, developer portals, workflow orchestration, Azure, APIs, integrations, containers, scripting, CMDB, asset management, discovery, and dependency mapping.
Experience in regulated environments with audit readiness, operational controls, change governance, evidence retention, and compliance expectations.
Education, Experience, and Certifications
Bachelor s degree in Information Technology, Computer Science, Engineering, or related field; equivalent experience may be considered.
ITIL Foundation preferred; cloud, observability, automation, or platform-specific certifications are a plus.
Success Measures / KPIs
Availability, uptime, performance, capacity health, and reliability trends.
Change/release success, patch currency, upgrade readiness, and rollback effectiveness.
Reduction in incidents, repeat issues, alert noise, MTTD, MTTR, and manual toil.
Accuracy of agent inventory, dependencies, lifecycle, licensing, documentation, and audit evidence."
Desirable Skills:
Keyword:
Skills: Digital : Splunk~AI and Automation~Digital : Site Reliability Engineering (SRE)
Experience Required: 6-8
Appreciate your quick response and please feel free to reach me out for any query you may have.
Thanks