Senior Site Reliability Engineer

Gruve

Redwood City (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

A software services startup based in Redwood City, California, is looking for a Site Reliability Engineer (SRE) to lead architectural improvements across IT and infrastructure. The successful candidate will have 6-9 years of experience, particularly with Kubernetes, as well as a strong coding background. This full-time, onsite role includes mentoring engineers and driving automation. Join this innovative team in a collaborative environment where your contributions can make a real impact.

Qualifications

  • 6-9 years of SRE or platform engineering experience.
  • Expert understanding of Kubernetes operations.
  • Strong coding background in Python, Go, or Java.
  • Deep knowledge of observability tools like Prometheus and Grafana.

Responsibilities

  • Lead reliability strategy and architectural improvements.
  • Mentor junior and mid‑level SREs.
  • Drive automation, IaC, and reliability tooling.

Skills

Kubernetes operations
Cloud platform experience (AWS/GCP/Azure)
Advanced networking
Security fundamentals
Python
Go
Java
Prometheus
Grafana
ELK / Fluentd

Job description

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises into AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well‑funded early‑stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

About the Role

This role leads reliability strategy and architectural improvements across infrastructure, GPU systems, observability, ML Ops and IT Ops. Mentor engineers, manage high‑severity incidents, and drive SLO governance. You will work with other SRE engineers to set up, maintain, and troubleshoot the stack from bare metal through the application layer.

Key Responsibilities
  • Architect reliability improvements across Kubernetes, GPU infrastructure, ML Ops, networking, and monitoring.
  • Lead incident management, blameless post‑mortems, and error‑budget policies.
  • Drive automation, IaC, and reliability tooling at scale.
  • Oversee metrics, logs, tracing, and dashboards; ensure actionable alerting.
  • Integrate GPU operators/exporters and model lifecycle workflows for inference platforms.
  • Mentor junior and mid‑level SREs and guide cross‑team initiatives.
Basic Qualifications
  • 6–9 years of SRE or platform engineering experience.
  • Expert Kubernetes operations and cloud platform experience (AWS/GCP/Azure).
  • Advanced networking and security fundamentals.
  • Strong coding background (Python, Go, or Java).
  • Deep observability knowledge (Prometheus, Grafana, ELK / Fluentd).
Preferred Qualifications
  • GPU reference architecture expertise and performance tuning.
  • Experience with chaos engineering, capacity planning, and multi‑region design.

This is an onsite, full‑time position with Gruve. The role is open at our Redwood City, California, and Edison, New Jersey offices.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Incident Response Investigator
Senior Incident Response Investigator

Gruve • Redwood City (CA)

On-site
USD 160,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Automotive Full-Stack Software Engineer
Automotive Full-Stack Software Engineer

Gruve • Palo Alto (CA)

On-site
USD 150,000 - 175,000
DevOps Director
DevOps Director

Gruve • United States

On-site
USD 200,000 - 230,000
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
DevOps Director
DevOps Director

Gruve • San Francisco (CA)

On-site
USD 200,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2