Senior/Lead Site Reliability Engineer Observability

Tata Consultancy Services

San Jose (CA)

On-site

USD 94,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Tata Consultancy Services in California, San Jose, seeks a senior Site Reliability/Platform Engineer with deep observability expertise to design, deploy, and operate enterprise-grade platforms. You will build and maintain Splunk Enterprise/Splunk Cloud, Elasticsearch clusters, and Grafana dashboards, while scaling monitoring with Tempo/OpenTelemetry and Kafka.

You will automate infrastructure with Terraform, manage Kubernetes-based ecosystems, and collaborate with security/compliance teams to

Qualifications

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
  • Experience supporting FedRAMP or regulated environments.

Responsibilities

  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.

Skills

Splunk
Elasticsearch/ELK
Prometheus
Grafana
Grafana Tempo
OpenTelemetry
Kafka
Terraform
Python
Go
Ruby
Bash
Kubernetes
AWS/Azure/GCP
Ansible
Consul
CI/CD pipelines
Service mesh technologies
FedRAMP

Tools

Terraform
Kubernetes
Docker
Linux
Kibana
Grafana
Tempo

Job description

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
  • Experience supporting FedRAMP or regulated environments.
Job Description
  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
  • Experience supporting FedRAMP or regulated environments.
Technology Stack

Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

Roles & Responsibilities
  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.
Nice To Have Skills
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.

In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person be a U.S. Citizen, a U.S. Permanent Resident (i.e., a “Green Card Holder”), or a Political Asylee or Refugee.

Salary Range: $94,000 - $130,000 a year

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Splunk & Observability Engineer
Senior Splunk & Observability Engineer

System One • Columbia (SC)

On-site
USD 150,000 - 185,000
Senior Splunk & Observability Engineer
Senior Splunk & Observability Engineer

System One • Birmingham (AL)

On-site
USD 120,000 - 180,000
Senior Splunk Infrastructure Engineer
Senior Splunk Infrastructure Engineer

Summit Tech Partners • Washington, Tacoma (WA)

Hybrid
USD 140,000 - 200,000
Splunk Administrator/Engineer
Splunk Administrator/Engineer

Resolution Technologies, Inc. • Georgia

On-site
USD 80,000 - 110,000
Splunk Subject Matter Expert (SME) & Enterprise Monitoring Engineer
Splunk Subject Matter Expert (SME) & Enterprise Monitoring Engineer

Empower Professionals Inc - Talent & IT Services • Frisco (TX)

Hybrid
USD 120,000 - 150,000
Splunk Observability Engineer
Splunk Observability Engineer

System One • Birmingham (AL)

On-site
USD 90,000 - 150,000
Splunk Observability Engineer
Splunk Observability Engineer

System One • Columbia (SC)

On-site
USD 120,000 - 150,000
Splunk Observability Engineer
Splunk Observability Engineer

System One • Lafayette (LA)

On-site
USD 100,000 - 130,000
Senior Engineer
Senior Engineer

Hobbsnews • Charlotte (NC)

On-site
USD 122,000 - 200,000
Benefits eligible
Paid time off
Annual discretionary award
Observability Operations Engineer
Observability Operations Engineer

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 100,000 - 120,000