Incident Manager

AssetMark, Inc.

  • Charlotte, NC
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Web Services (AWS)unmatched
    • Atlassian JIRAunmatched
    • Automationunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Consultingunmatched
    • Continuous Improvementunmatched
    • Cross-Functionalunmatched
    • Customer Support/Serviceunmatched
    • Firefightingunmatched
    • GCP (Good Clinical Practices)unmatched
    • High Availabilityunmatched
    • ITIL (IT Infrastructure Library)unmatched
    • Identify Issuesunmatched
    • Incident Managementunmatched
    • Leadershipunmatched
    • Machine Toolunmatched
    • Microsoft Windows Azureunmatched
    • Multiplatform/Cross-Platformunmatched
    • On Callunmatched
    • Problem Solving Skillsunmatched
    • Process Improvementunmatched
    • Production Systemsunmatched
    • Reliability Engineeringunmatched
    • Root Cause Analysisunmatched
    • Sales Closing Skillsunmatched
    • ServiceNowunmatched
    • Systems Reliabilityunmatched
    • User Interface/Experience (UI/UX)unmatched

    Description

    Job Description:

    AssetMark is a leading strategic provider of innovative investment and consulting solutions serving independent financial advisors. We provide investment, relationship, and practice management solutions that advisors use in helping clients achieve wealth, independence, and purpose.

    The Job/What Youll Do:

    We are looking for an experienced Incident Manager to own the end-to-end lifecycle of production issues across our technology platforms and services.

    This role goes beyond traditional incident coordination. Incident Managers are hands-on operators responsible for driving rapid service restoration, resolving issues directly whenever possible, and eliminating recurring problems at their source.

    You will work across the full technology stack-partnering with engineering, infrastructure, and operations teams-to ensure reliable system performance and a high-quality user experience.

    This is a high-visibility role that requires strong technical judgment, clear communication under pressure, and a bias toward action. You will play a critical role in improving system reliability while helping teams spend less time firefighting and more time building.

    This role participates in a 24/7 operating model, including on-call responsibilities.

    We can only consider candidates for this position who are able to accommodate a hybrid work schedule and are close to our Charlotte, NC office.

    Responsibilities:

    End-to-End Problem Management (Sev1-Sev5)

    Own production issues from detection through full resolution

    Quickly assess impact and assign severity (Sev1-Sev5)

    Lead triage, investigation, and resolution efforts

    Maintain clear ownership throughout the lifecycle, regardless of which teams are involved

    Drive fast, effective restoration of service

    Resolve More, Closer to the Team

    Directly investigate and resolve issues whenever possible

    Partner closely with operations and reliability teams to resolve issues without unnecessary escalation

    Reduce dependency on engineering teams for repeat or well-understood problems

    Build reusable knowledge and patterns to improve team self-sufficiency

    Root Cause Analysis & Prevention

    Perform and/or lead root cause analysis (RCA)

    Identify recurring patterns and systemic weaknesses

    Drive fixes that prevent entire classes of issues from recurring

    Ensure issues are fully resolved-not just temporarily mitigated

    Incident Leadership & Communication

    Lead real-time response for high-impact production issues

    Coordinate cross-functional teams with clarity and urgency

    Communicate clearly with stakeholders, including leadership, during active incidents

    Provide structured updates on impact, progress, and next steps

    Process, Tooling & Continuous Improvement

    Improve incident management processes, workflows, and operating models

    Build and maintain runbooks and response procedures

    Identify opportunities for automation and better monitoring

    Ensure high-quality documentation and knowledge sharing

    What You Bring

    Required Experience & Skills

    5+ years of experience in incident management, site reliability engineering (SRE), production operations, or similar roles

    Proven ability to lead and resolve production issues under pressure

    Strong technical breadth across systems, applications, and infrastructure

    Ability to diagnose and troubleshoot issues directly, not just coordinate response

    Excellent communication skills-clear, concise, and composed under pressure

    Strong sense of ownership and accountability

    Analytical mindset with strong problem-solving skills

    Preferred Qualifications

    Experience in high-availability, large-scale production environments

    Familiarity with tools such as ServiceNow, Jira Service Management, or PagerDuty

    Experience with cloud platforms (AWS, Azure, or GCP)

    Familiarity with monitoring and observability tools

    Knowledge of ITIL frameworks (helpful, but not required)

    How We Measure Success

    Success in this role is defined by outcomes:

    Faster time to restore service (MTTR)

    More issues resolved directly within the incident management / operations function

    Reduction in high-severity issues (Sev1 / Sev2)

    Fewer recurring issues due to strong root cause resolutio

    Numbers & Facts

    LocationCharlotte, NC

    Similar Jobs