Senior SRE: Kubernetes, GPU Infra & ML Ops Leader

Gruve

Redwood City (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A software services startup based in Redwood City, California, is looking for a Site Reliability Engineer (SRE) to lead architectural improvements across IT and infrastructure. The successful candidate will have 6-9 years of experience, particularly with Kubernetes, as well as a strong coding background. This full-time, onsite role includes mentoring engineers and driving automation. Join this innovative team in a collaborative environment where your contributions can make a real impact.

Qualifications

  • 6-9 years of SRE or platform engineering experience.
  • Expert understanding of Kubernetes operations.
  • Strong coding background in Python, Go, or Java.
  • Deep knowledge of observability tools like Prometheus and Grafana.

Responsibilities

  • Lead reliability strategy and architectural improvements.
  • Mentor junior and mid‑level SREs.
  • Drive automation, IaC, and reliability tooling.

Skills

Kubernetes operations
Cloud platform experience (AWS/GCP/Azure)
Advanced networking
Security fundamentals
Python
Go
Java
Prometheus
Grafana
ELK / Fluentd

Job description

A software services startup based in Redwood City, California, is looking for a Site Reliability Engineer (SRE) to lead architectural improvements across IT and infrastructure. The successful candidate will have 6-9 years of experience, particularly with Kubernetes, as well as a strong coding background. This full-time, onsite role includes mentoring engineers and driving automation. Join this innovative team in a collaborative environment where your contributions can make a real impact.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Access to BetterUp coaching
Competitive compensation plan
Medical, dental, and vision insurance
+3
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Site Reliability Engineer - Kubernetes & Cloud
Site Reliability Engineer - Kubernetes & Cloud

Hydrolix • United States

On-site
USD 110,000 - 150,000
Senior SRE Engineer – Cloud, Kubernetes & Automation
Senior SRE Engineer – Cloud, Kubernetes & Automation

VBeyond Corporation • Dallas (TX)

On-site
USD 100,000 - 140,000
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Competitive compensation with equity
100% medical, dental, and vision coverage
Generous PTO including Winter Break
+2
Senior SRE Lead – Cloud-Native, AI/ML, Kubernetes, Austin
Senior SRE Lead – Cloud-Native, AI/ML, Kubernetes, Austin

Compunnel, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
SRE Engineer: Cloud, Kubernetes & CI/CD Reliability
SRE Engineer: Cloud, Kubernetes & CI/CD Reliability

OPPO • Palo Alto (CA)

On-site
USD 100,000 - 200,000
Senior SRE - Cloud Reliability & Kubernetes (Onsite Irvine)
Senior SRE - Cloud Reliability & Kubernetes (Onsite Irvine)

Ledgent Technology • Irvine (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer: Cloud, Kubernetes & CI/CD
Senior Site Reliability Engineer: Cloud, Kubernetes & CI/CD

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000