Senior/Lead Site Reliability Engineer Observability

Tata Consultancy Services Ltd
  • San Jose, CA
  • $94,000–$130,000 Per Year
6 days ago

Job Description

Role: Functional Consultant

Required Skills: DevOps

Experience: 8 - 10 Years

Function: TECHNOLOGY

Apply By: 2026-10-10

Must Have Technical/Functional Skills:

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
  • Experience supporting FedRAMP or regulated environments.

Technology Stack

Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

Roles & Responsibilities:

  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.

Nice to have skills:

  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.

In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person be a U.S. Citizen, a U.S. Permanent Resident (i.e., a "Green Card Holder"), or a Political Asylee or Refugee.

Salary Range: $94,000 - $130,000 a year

#LI-CM2

Numbers & Facts

LocationSan Jose, CA
Salary$94,000–$130,000 Per Year

Skills

  • Amazon Web Services (AWS)unmatched
  • Ansibleunmatched
  • Apache Kafkaunmatched
  • Bash Scriptingunmatched
  • Cloud Computingunmatched
  • Configuration Managementunmatched
  • Consultingunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • DevOpsunmatched
  • Elasticsearchunmatched
  • Forwarderunmatched
  • GCP (Good Clinical Practices)unmatched
  • Go Programming Language (Golang)unmatched
  • Instrumentationunmatched
  • Metricsunmatched
  • Microsoft Windows Azureunmatched
  • Python Programming/Scripting Languageunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Retention Programsunmatched
  • Rubyunmatched
  • Splunkunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder