Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai

San Francisco (CA)

On-site

USD 170,000 - 290,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Luma AI is seeking a hands-on SRE/Infrastructure Engineer to design and operate our GPU-driven AI infrastructure across on-prem and multi-cloud environments. You will own end-to-end reliability, performance, and security in a fast-paced startup setting.

From Linux performance tuning to building automation with Python/Go, you will collaborate with hardware vendors like NVIDIA and drive scalability for massive distributed training workloads.

Qualifications

  • 5+ years as an SRE/production infra engineer in large-scale environments.
  • Deep Linux, containerized systems, and low-level performance debugging.

Responsibilities

  • Architect for Reliability and Scale across GPU infrastructure.
  • Own multi-cloud GPU clusters (AWS & OCI) with high availability.
  • Drive security and compliance (SOC 2, ISO) measures in startup infra.
  • Perform deep Linux performance tuning at OS/kernel level.
  • Build automation tools in Python, Go, or Bash for infra management.
  • Debug complex hardware/software failures with hardware vendors.

Skills

Linux
Terraform
Airflow
Ray
Kubernetes
Networking
Python
Go
Bash

Tools

AWS
OCI
DCGM/ROCm
InfiniBand/RDMA

Job description

Luma AI is seeking a hands-on SRE/Infrastructure Engineer to design and operate our GPU-driven AI infrastructure across on-prem and multi-cloud environments. You will own end-to-end reliability, performance, and security in a fast-paced startup setting.

From Linux performance tuning to building automation with Python/Go, you will collaborate with hardware vendors like NVIDIA and drive scalability for massive distributed training workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Staff AI Infrastructure Engineer - Frontier GPU Scale
Staff AI Infrastructure Engineer - Frontier GPU Scale

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 360,000
Senior Cloud Platform Engineer — AI GPU Infra
Senior Cloud Platform Engineer — AI GPU Infra

Lambda • San Francisco (CA)

On-site
USD 180,000 - 320,000
Health, dental, and vision
401k with 2% match
Wellness stipend
+1
Senior SRE: AI Infra on-Site in SF, GPU & Cloud
Senior SRE: AI Infra on-Site in SF, GPU & Cloud

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE: GPU Fleet Orchestration & Auto-Scaling
Senior SRE: GPU Fleet Orchestration & Auto-Scaling

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits