Senior SRE: AI Compute, Kubernetes & Observability

Justjoin

United States

Remote

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
Financial planning
Family benefits
Work-life balance
Time for other pursuits

Job summary

Justjoin seeks a Senior SRE to own reliability workstreams for a serverless inference platform, build automation, and drive architecture and operational decisions. You will partner with product engineering to scale GPU infrastructure and AI workloads using Kubernetes at scale.

You will lead observability efforts, implement automation to reduce toil, and contribute to incident management with runbooks and blameless post-mortems.

Qualifications

  • Expertise in SRE, infra, or platform engineering for large-scale distributed systems.
  • Kubernetes and large-scale containerization experience.
  • Define SLOs and use observability tools (Prometheus, Grafana, tracing).
  • Proficiency in Python or Go for automation and IaC (Terraform).
  • Interest in AI/ML infra, model serving, or GPU workloads.
  • Independent problem solver with accountability.
  • Collaborate with teams unfamiliar with SRE practices.

Responsibilities

  • Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • Integrating AI workloads into our existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
  • Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
  • Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Skills

SRE practices
Automation
Observability
On-call experience

Tools

Kubernetes
Terraform
Prometheus
Grafana
Distributed tracing

Job description

Justjoin seeks a Senior SRE to own reliability workstreams for a serverless inference platform, build automation, and drive architecture and operational decisions. You will partner with product engineering to scale GPU infrastructure and AI workloads using Kubernetes at scale.

You will lead observability efforts, implement automation to reduce toil, and contribute to incident management with runbooks and blameless post-mortems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Justjoin • United States

Remote
USD 140,000 - 190,000
Health benefits
Financial planning
Family benefits
+2
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior SRE, AI Platform: Scale Reliability & Kubernetes
Senior SRE, AI Platform: Scale Reliability & Kubernetes

United States Digital Space LLC • Paris (TX)

On-site
USD 80,000 - 111,000
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1
Senior SRE: AI-First Platform & Observability Lead (Remote)
Senior SRE: AI-First Platform & Observability Lead (Remote)

Bot Jobs • Myrtle Point (OR)

Remote
USD 170,000 - 210,000
In-person connection offsites
Tech & learning stipend
Remote by design
+2
Senior SRE: AI Infra, On-Call, Kubernetes & Ceph
Senior SRE: AI Infra, On-Call, Kubernetes & Ceph

Engg • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000