Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma

Redwood City (CA)

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma, based in California, seeks a Senior SRE to own the GPU infrastructure that powers our research and product workloads across on-prem, AWS, and OCI. You will keep training and inference clusters reliable and fast, collaborating on scale across multiple environments.

This hands-on role is Linux-focused and near the metal—debugging GPU, networking, and kernel-level issues, including direct work with NVIDIA. If you want a narrowly scoped ops job, this isn’t it.

Qualifications

  • 5+ years as an SRE, production, or infra engineer in a fast-paced, large-scale environment.
  • Deep Linux expertise, containerized systems, and low-level performance debugging.
  • Experience with AWS or OCI and high-performance networking (InfiniBand/RDMA).
  • Security and compliance familiarity (SOC 2 / ISO).

Responsibilities

  • Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI.
  • Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply, at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
  • Serve as the final escalation for the hardest GPU, networking, and system failures, working with vendors like NVIDIA.
  • Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.

Skills

Linux kernel debugging
GPU infrastructure
Terraform
Airflow
Ray
AWS
OCI
InfiniBand/RDMA
SOC 2 / ISO security
Kubernetes
GPU tooling (DCGM/ROCm)
Large-scale GPU clusters
Kubernetes / Ray orchestration
Data pipelines & infra

Tools

Terraform
Airflow
Ray
Kubernetes

Job description

Luma, based in California, seeks a Senior SRE to own the GPU infrastructure that powers our research and product workloads across on-prem, AWS, and OCI. You will keep training and inference clusters reliable and fast, collaborating on scale across multiple environments.

This hands-on role is Linux-focused and near the metal—debugging GPU, networking, and kernel-level issues, including direct work with NVIDIA. If you want a narrowly scoped ops job, this isn’t it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior SRE: FinOps-Driven Infra, GPU & Scale
Senior SRE: FinOps-Driven Infra, GPU & Scale

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Senior GPU Platform SRE | Kubernetes, Slurm, BCM | Equity
Senior GPU Platform SRE | Kubernetes, Slurm, BCM | Equity

Visa Hunt • United States

On-site
USD 168,000 - 334,000
Equity
Benefits
Remote Senior SRE: GPU Cloud, Automation & Reliability
Remote Senior SRE: GPU Cloud, Automation & Reliability

asobbi • United States

On-site
USD 150,000 - 190,000
Equity
Bonus
Benefits
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1