Senior/Staff Cloud Reliability Engineer

ThoughtSpot

Mountain View (CA)

On-site

USD 180,000 - 240,000

Full time

13 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ThoughtSpot is seeking a Staff Site Reliability Engineer to own the availability, reliability, security, and efficiency of our multi-cloud SaaS platform on AWS and GCP. You will drive automation, AI-augmented operations, and incident response for large-scale Kubernetes-based control and data planes.

Required are 6+ years of SaaS ops, Go/Python proficiency, IaC expertise with Terraform/Ansible, and strong cloud security know-how.

Qualifications

  • B.Tech. degree in Computer Science or equivalent.
  • 6+ years of Enterprise SaaS Ops experience.
  • Strong proficiency in Go and Python; IaC tools like Terraform/Ansible.
  • Experience with cloud security and cloud networking.
  • Experience with AI Ops tools and Agentic LLMs.

Responsibilities

  • Operate a high-scale multi-cloud SaaS platform (AWS, GCP) for reliability and uptime.
  • Embed AI/Agentic workflows into SRE practices for anomaly detection and automated triage.
  • Drive capacity planning and autoscaling to support SaaS growth.
  • Architect and operate Kubernetes controller frameworks and data/control planes.

Skills

Go
Python
Kubernetes
IaC
Cloud security
AI Ops
PostgreSQL

Education

B.Tech. in CS

Tools

Terraform
Ansible
GitOps
LLM/AI Ops tools

Job description

We are seeking a Staff Site Reliability Engineer with deep enterprise SaaS operations expertise to own the availability, reliability, security, and efficiency of our Multi-Cloud (AWS, GCP) production SaaS platform. The ideal candidate brings hands‑on experience running highly available, large-scale Kubernetes‑based control and data planes, a strong bias toward automation and AI‑augmented operations, and a proven track record in production security, capacity management, and cloud‑native data infrastructure.

Responsibilities
  • Operate a high-scale, multi-cloud (AWS, GCP) SaaS platform — ensuring reliability, performance, and uptime for business‑critical production workloads.
  • Embed AI and Agentic workflows into SRE practice: leverage AI Ops platforms and LLM‑powered autonomous agents for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
  • Drive capacity planning and scaling operations — proactively model growth, right‑size infrastructure, and implement horizontal/vertical autoscaling strategies to support SaaS growth without reliability regression.
  • Architect and operate Kubernetes controller frameworks governing both control plane and data plane services; define and enforce operational standards for cluster lifecycle, workload scheduling, autoscaling, and failover.
  • Own operations of high‑scale cloud‑native databases and data infrastructure: PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache (Redis/Memcached) on AWS and GCP — including performance tuning, backup/recovery, and incident response.
  • Lead incident response and blameless post‑mortems for P0/P1 events; drive root cause analysis to permanent resolution and prevention — eliminating repeat incidents through systemic fixes, not workarounds.
  • Define and enforce a culture of automation‑first: identify and eliminate toil through self‑healing systems, automated remediation pipelines, and infrastructure‑as‑code (Terraform, Helm, GitOps).
  • Participate in on‑call rotations for critical cloud infrastructure; serve as a senior escalation point and incident commander during high‑severity events.
  • Achieve quantifiable SaaS operational Excellence measured by related SLI/SLO/SLA
Required skills/qualifications
  • B.Tech. degree in Computer Science or equivalent.
  • At least 6+ years of Enterprise SaaS Ops experience
  • Strong proficiency in programming, particularly with Go and Python, and experience with Infrastructure as Code (IaC) tools like Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking
  • Experience with AI Ops tools, Agentic LLM.
  • Experience/ Knowledge in Cloud Services, Kubernetes, Cloud Databases like Postgres/RDS/MySQL/DynamoDB, Elastic, Kafka, and Microservice architecture is a bonus.
  • Experience in implementing and operating enterprise‑grade observability ( metrics, logs, tracing), alerting stack in a Cloud SaaS environment
  • Strong debugging and problem‑solving skills (network, systems, database, and application).
  • Advanced professional certifications from Cloud Providers ( AWS, Azure, GCP) in domains like K8s, Solution architecture, networking, and databases are a bonus.
  • Full Stack Architecture/Development Experience is a bonus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 190,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Software Engineer – Cloud Platform & Operations
Senior Software Engineer – Cloud Platform & Operations

TalentCloud Group • Indiana

Hybrid
USD 100,000 - 130,000
Competitive salary
Equity
Flexible working hours
+1
AI-Driven Multi-Cloud SRE Leader
AI-Driven Multi-Cloud SRE Leader

ThoughtSpot • Mountain View (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • San Francisco (CA)

On-site
USD 140,000 - 200,000