Principal Engineer

Mindlance

  • ISELIN, NJ
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Analysis Skillsunmatched
    • Automationunmatched
    • Backlog Prioritizationunmatched
    • Best Practicesunmatched
    • Budgetingunmatched
    • Capacity and Performance Managementunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computer Networksunmatched
    • Continuous Improvementunmatched
    • Detail Orientedunmatched
    • Distributed Computingunmatched
    • Enterprise Applicationsunmatched
    • Environmental Impactunmatched
    • High Availabilityunmatched
    • High Reliabilityunmatched
    • IT Service Management (ITSM)unmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Leadershipunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Middlewareunmatched
    • Multiplatform/Cross-Platformunmatched
    • Multitaskingunmatched
    • Network Administration/Managementunmatched
    • Performance Engineeringunmatched
    • Performance Managementunmatched
    • Performance Metricsunmatched
    • Problem Solving Skillsunmatched
    • Process Improvementunmatched
    • Production Supportunmatched
    • Production Systemsunmatched
    • Purchasing/Procurementunmatched
    • Registered Training Organisation (RTO)unmatched
    • Reliability Engineeringunmatched
    • Reporting Dashboardsunmatched
    • Software Administrationunmatched
    • System Operationsunmatched
    • Systems Reliabilityunmatched
    • Team Playerunmatched
    • Technical Leadershipunmatched
    • Trend Analysisunmatched
    • Usage Analysisunmatched
    • User Interface/Experience (UI/UX)unmatched

    Description

    Objective: Improve scalability & reliability of existing systems. Drive retroactive mitigation of technical debt for critical systems

    Skills and Focus of PREs

    • O bservability
      • Alerts & Dashboards focused on user experience
      • Focus: Metrics, Logs, Traces
    • C apacity
      • Continuous Trend & Usage analysis
      • Thresholds for new equipment purchases with lead time
    • P erformance
      • Measure performance from user perspective
      • Focus: Response time, Error rate, Latency
    • R esiliency
      • High Availability & Fault Tolerance (within & across DC)
      • Focus: RTO should match business expectations


    In this role, you will:
    • Act as a Platform Reliability Engineering (PRE) expert, providing deep technical leadership in one core domain (Database, Cloud, Network, Compute/Storage, Middleware, or Application Support), while partnering across broader platform teams.
    • Lead analysis and remediation of systemic reliability issues and complex production problems, translating recurring incidents into long-term engineering solutions.
    • Design, influence, and implement scalable, highly reliable systems using SRE principles including SLI/SLOs, error budgets, and automation-first approaches.
    • Drive observability standards across platforms, with a strong focus on metrics, logs, traces, and user-experience-based alerting and dashboards.
    • Partner with application, infrastructure, cloud, and support teams to improve availability, performance, capacity, and resiliency of both new and existing systems.
    • Lead or contribute to blameless post-mortems, ensuring actionable outcomes and sustained reduction of repeat incidents.
    • Translate advanced technical knowledge and enterprise context into clear guidance for senior leadership on reliability risks, priorities, and investment areas.
    • Mentor and guide engineers and support staff on reliability best practices, operational standards, and automation opportunities.

    Required Qualifications:
    • 10+ years of experience in Systems Operations, SRE, Platform Engineering, or Production Support, with deep expertise in at least one of the following domains:
      • Database platforms
      • Cloud platforms
      • Network infrastructure
      • Compute / Storage platforms
      • Middleware platforms
      • Enterprise Application Support
    • Strong hands-on experience applying SRE concepts such as SLI/SLO definition, error budgets, reliability metrics, and incident-driven engineering improvements.
    • Proven experience diagnosing and resolving complex, large-scale production issues across distributed systems.
    • Solid understanding of observability tools and practices covering monitoring, alerting, logging, and tracing.
    • Experience driving automation and self-service to reduce operational toil and manual interventions.
    • Strong communication skills with the ability to influence engineers, partners, and senior leaders.

    Desired Qualifications:
    • Exposure to capacity management, performance engineering, and resiliency design (HA, fault tolerance, RTO/RPO).
    • Experience working in environments with hybrid platforms (on-prem + cloud) and complex enterprise dependencies.
    • Ability to drive technical debt remediation for critical legacy systems using structured, prioritized backlogs.
    • Familiarity with IT service management, incident/problem management, and continuous improvement frameworks.
    • Experience mentoring or leading senior engineers in reliability or operations-focused roles.

    Job Expectations:
    • Strong collaboration and partnering skills across platform, application, and support teams.
    • Ability to manage multiple priorities in a fast-paced, high-impact production environment.
    • Consistent delivery of high-quality reliability outcomes within expected timelines.
    • High attention to detail, data-driven problem-solving, and operational rigor.
    • Prior project or initiative leadership experience is highly desirable.

    Benefits:

    • Health insurance
    • 401(k)

    Numbers & Facts

    LocationISELIN, NJ

    Similar Jobs

    See more jobs