Site Reliability Engineer

Artech LLC
  • Charlotte, NC
  • Quick Apply
30+ days ago

Job Description

Location:

Charlotte, NC

Salary Range:

Competitive, based on experience

Introduction

We are seeking a highly skilled and experienced professional to join our team as a Site Reliability Engineer III. This role involves collaborating with cross-functional teams to ensure the reliability and performance of critical systems. The ideal candidate will have a strong technical background and a passion for improving system efficiency and reliability.

Required Skills & Qualifications

  • Experience with APM tools such as DynaTrace.
  • Familiarity with cloud platforms like AWS, Azure, or Google Cloud.
  • Knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools.
  • Knowledge in monitoring and logging tools such as Prometheus, Grafana, ELK stack, or Splunk.
  • Prior experience designing and supporting Enterprise applications.

Preferred Skills & Qualifications

  • Expertise in APM tools, i.e., DynaTrace.
  • Strong problem-solving and troubleshooting skills, with the ability to analyze and resolve complex technical issues.
  • Excellent communication and collaboration skills to work effectively with cross-functional teams.
  • Understanding of networking principles and protocols (TCP/IP, HTTP, DNS, etc.).
  • Strong attention to detail and ability to work in a fast-paced, dynamic environment.
  • Prior work experience at client or in client's industry.
  • Applicants must be able to work directly for Artech on W2.
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
  • Strong knowledge of Linux/Unix systems and command line tools.
  • Proficiency in scripting languages such as Python, Shell, or Perl.

Day-to-Day Responsibilities

  • Collaborate with cross-functional teams to define and establish service level objectives (SLOs) and service level agreements (SLAs) for critical systems.
  • Monitor systems and applications, proactively identifying and resolving any performance bottlenecks or availability issues.
  • Develop and maintain monitoring tools, alerts, and dashboards to provide visibility into system health and performance.
  • Conduct post-incident analyses to identify root causes and implement preventive measures to avoid future incidents.
  • Automate repetitive tasks and processes to improve efficiency and reduce manual intervention.
  • Create and maintain documentation for system architecture, configuration, and troubleshooting procedures.
  • Collaborate with development teams to implement and deploy new features and enhancements, ensuring they meet reliability and performance standards.

Company Benefits & Culture

  • Comprehensive health, dental, and vision insurance.
  • Flexible work schedule and remote work options.
  • Opportunities for professional growth and development.

For immediate consideration please click APPLY to begin the screening process.

Numbers & Facts

LocationCharlotte, NC

Skills

  • Amazon Web Services (AWS)unmatched
  • Analysis Skillsunmatched
  • Cloud Computingunmatched
  • Command Lineunmatched
  • Communication Skillsunmatched
  • Computer Scienceunmatched
  • Cross-Functionalunmatched
  • DNS (Domain Name System)unmatched
  • Dental Insuranceunmatched
  • Detail Orientedunmatched
  • Dockerunmatched
  • Documentationunmatched
  • Enterprise Applicationsunmatched
  • HTTP (HyperText Transport Protocol)unmatched
  • Identify Issuesunmatched
  • Linux Operating Systemunmatched
  • Microsoft Windows Azureunmatched
  • Network Protocolsunmatched
  • Perl Programming Languageunmatched
  • Problem Solving Skillsunmatched
  • Process Improvementunmatched
  • Python Programming/Scripting Languageunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Root Cause Analysisunmatched
  • Scripting (Scripting Languages)unmatched
  • Service Level Agreement (SLA)unmatched
  • Software Administrationunmatched
  • Splunkunmatched
  • System Architectureunmatched
  • Systems Administration/Managementunmatched
  • Systems Maintenanceunmatched
  • Systems Reliabilityunmatched
  • TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
  • Team Playerunmatched
  • Testingunmatched
  • Unix Operating Systemsunmatched
  • Unix Shell Programmingunmatched
  • Vision Planunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder