Senior GPU Infra Reliability Engineer - Remote

Luma AI

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI is seeking a Senior SRE to own the GPU infrastructure used for training and inference across on-prem, AWS, and OCI. You’ll be hands-on, close-to-the-metal, tackling GPU, networking, and kernel-level failures and driving reliability at scale.

You’ll tune Linux performance, build automation in Python/Go/Bash, and participate in security and compliance efforts while partnering with NVIDIA and other vendors in a fast, unstructured environment.

Qualifications

  • 5+ years as SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep Linux expertise, containerized systems, and kernel/OS-level performance debugging.
  • Hands-on experience with cloud platforms (AWS or OCI) and high-speed networking.

Responsibilities

  • Take end-to-end ownership of production GPU clusters for training and inference across AWS/OCI.
  • Join re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure.
  • Serve as final escalation for GPU, networking (InfiniBand/RDMA), and system failures with vendors like NVIDIA.
  • Support security certifications (SOC 2 Type 1/2, ISO) with strong infrastructure practices.

Skills

Linux expertise
Containerized systems
Performance debugging
Networking experience

Tools

Terraform
Airflow
Ray
AWS
OCI
Kubernetes

Job description

Luma AI is seeking a Senior SRE to own the GPU infrastructure used for training and inference across on-prem, AWS, and OCI. You’ll be hands-on, close-to-the-metal, tackling GPU, networking, and kernel-level failures and driving reliability at scale.

You’ll tune Linux performance, build automation in Python/Go/Bash, and participate in security and compliance efforts while partnering with NVIDIA and other vendors in a fast, unstructured environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior SRE, AI Infrastructure & GPU Fleet Reliability
Senior SRE, AI Infrastructure & GPU Fleet Reliability

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior Backend Engineer, GPU Cloud Infra & Kubernetes
Senior Backend Engineer, GPU Cloud Infra & Kubernetes

Socket.dev • New York (NY)

Hybrid
USD 180,000 - 250,000
Health insurance
Equity
401(k) matching
+6
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000