Senior Site Reliability Engineer - AI platform services

Applab Systems Inc
  • Santa Clara, CA
  • Instant Apply
30+ days ago

Job Description

Role: - Senior Site Reliability Engineer - AI platform services
Location: - Santa Clara, CA(Onsite)
Duration- 1 Year

ENGAGEMENT SUMMARY

The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on
reliability engineering, incident response, service health, and operational automation. This role is best suited to a senior hands-on engineer who can improve availability while remaining effective in detailed production
troubleshooting.

WHAT THIS CANDIDATE WILL BE DOING

" Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure
components.
" Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network, scheduling, job orchestration, or dependency failures.
" Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.
" Analyze recurring failure patterns and convert manual operations into automation and preventive controls.
" Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and event correlation.
" Troubleshoot performance and availability issues affecting training jobs, inference services, internal
platforms, and support tooling.
" Partner with infrastructure and validation teams to improve production readiness and change safety.
" Drive operational reviews, readiness criteria, and resilience testing.

WHAT WE NEE D TO SEE

" 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.
" Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.
" Experience building observability, alerting, and response workflows in complex production environments.
" Ability to balance urgent operational response with medium-term reliability engineering improvements.
" Strong scripting and automation skills, with experience reducing toil through tooling.
" Experience participating in incident management, root cause analysis, and post-incident follow-through.
" Strong communication skill with the ability to summarize technical issues clearly for cross-functional teams.

PREFERRED EXPERIENCE

" Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.
" Familiarity with Prometheus, Grafana, ELK/OpenSearch, Loki, PagerDuty, and incident tooling.
" Experience defining error budgets and applying SRE practices in environments with heavy batch and service traffic.

Numbers & Facts

LocationSanta Clara, CA

Skills

  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Budgetingunmatched
  • Communication Skillsunmatched
  • Cross-Functionalunmatched
  • Distributed Computingunmatched
  • Event Correlationunmatched
  • Failure Analysisunmatched
  • Follow Throughunmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Linux Operating Systemunmatched
  • Machine Toolunmatched
  • Metricsunmatched
  • Operational Auditunmatched
  • Pattern Analysisunmatched
  • Production Systemsunmatched
  • Reliability Engineeringunmatched
  • Root Cause Analysisunmatched
  • Scripting (Scripting Languages)unmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder