Manager - SRE

Mphasis Ltd

  • TX
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Analysis Skillsunmatched
    • Ansibleunmatched
    • Automationunmatched
    • Automation Engineeringunmatched
    • Cloud Computingunmatched
    • Elasticsearchunmatched
    • Google Search Engineunmatched
    • Leadershipunmatched
    • Linux Operating Systemunmatched
    • Operational Supportunmatched
    • Production Supportunmatched
    • Python Programming/Scripting Languageunmatched
    • Reporting Dashboardsunmatched
    • Service Level Agreement (SLA)unmatched
    • Simulationunmatched
    • Software Administrationunmatched
    • Software Developmentunmatched
    • Splunkunmatched
    • Team Lead/Managerunmatched
    • Technical Operationsunmatched
    • Technical Writingunmatched
    • Unix Shell Programmingunmatched

    Description

    Job Description

    Role: Automation Lead

    Automation Lead - Leading Automation SRE

    Responsible to perform end-to-end Self-Healing automation solution to reduce manual effort and TOIL.

    Primary Skill - Python Ansible Observability SRE

    Secondary Skill - Shell Script Linux Monitoring tools - Splunk AppD Grafana ITRS etc.

    Automation Engineer

    12 years of experience in leading Automation SRE teams.

    Advanced working experience with two or more of the following:

    Unix Linux Windows Server Oracle MSSQL MongoDB

    Experience with Python Java Curl scripting or any other types of scripting.

    Experience with two or more of the following observability tools:

    AppDynamics Big Panda Elastic Search ELK Google Cloud Logging Grafana Prometheus Splunk Thousand Eyes.

    Experience with logging monitoring and event detection on Cloud or Distributed platforms.

    Experience working with one or more of the following:

    AutoSys CRON Windows Scheduler or other logical batch schedulers.

    Provides technical direction regarding monitoring and logging to less experienced staff or develops highly complex original solutions.

    Acts as an Expert technical resource for modeling simulation and analysis efforts.

    Experience creating and modifying technical documentation such as environment flow functional requirements non-functional requirements.

    Outstanding problem solving and analytical skills with ability to turn findings into strategic imperatives.

    Technical operations application support experience.

    Minimum 4-6 years of hands-on experience into SRE implementation of monitoring system development for application reliability using Splunk Grafana App Dynamics Big panda.

    Job Nature

    1. Collaborate with Production support team identify the existing manual activities and automate.
    2. Identify toil area where it can be automated to avoid manual intervention.
    3. Build Monitoring system and observability platform for more Stack traces and s and Dashboards.
    4. Ability to define SLA SLO and SLI and implement the same for better monitoring.
    5. Scalability reliability and observability are the primary goals for reduction of MTTD and MTTR.

    Numbers & Facts

    LocationTX

    Similar Jobs