Remote Senior SRE: GPU Cloud, Automation & Reliability

asobbi

United States

On-site

USD 150,000 - 190,000

Full time

7 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Bonus
Benefits

Job summary

asobbi is seeking a Senior Site Reliability / DevOps Engineer to own the reliability and operability of major GPU cloud components in a remote US role. You will lead incident response, build automation, and IaC across Kubernetes, Linux, storage and network layers, with strong ownership and autonomy in a small, senior team.

You’ll mentor peers, drive on-call processes, and help bring new GPU clusters online while using modern tooling and AI-assisted operations to reduce manual toil and improve

Qualifications

  • 5+ years of SRE/DevOps/production engineering or infrastructure operations.
  • Strong hands-on Linux experience in production environments.
  • Deep operational Kubernetes experience, including troubleshooting at scale.

Responsibilities

  • Own reliability and performance across major production platform components.
  • Lead technical response during production incidents and drive problems through to resolution.
  • Build automation and Infrastructure as Code to remove repetitive operational work.
  • Improve monitoring, alerting and observability across production infrastructure.
  • Debug complex Linux, networking, storage and performance issues.
  • Support the deployment and operational readiness of new infrastructure and GPU clusters.
  • Create practical runbooks and post-mortems that improve future operations.
  • Mentor engineers through reviews, pairing and incident response.

Skills

Linux in production
Kubernetes
Infrastructure as Code
Terraform
Ansible
On-call experience
Incident response
Networking fundamentals
GPU/HPC infrastructure

Tools

Terraform
Ansible
Prometheus
Grafana

Job description

asobbi is seeking a Senior Site Reliability / DevOps Engineer to own the reliability and operability of major GPU cloud components in a remote US role. You will lead incident response, build automation, and IaC across Kubernetes, Linux, storage and network layers, with strong ownership and autonomy in a small, senior team.

You’ll mentor peers, drive on-call processes, and help bring new GPU clusters online while using modern tooling and AI-assisted operations to reduce manual toil and improve

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Remote Senior Kubernetes Platform Engineer - AI & GPUs
Remote Senior Kubernetes Platform Engineer - AI & GPUs

Sira Consulting, an Inc 5000 company • United States

On-site
USD 140,000 - 210,000
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior SRE: Lead Reliable AI Platform & Mentor the Team
Senior SRE: Lead Reliable AI Platform & Mentor the Team

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace