Senior Site Reliability Engineer

lumalabs-ai

San Francisco (CA)

On-site

USD 170,000 - 290,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma AI is seeking a hands-on SRE/Infrastructure Engineer to design and operate our GPU-driven AI infrastructure across on-prem and multi-cloud environments. You will own end-to-end reliability, performance, and security in a fast-paced startup setting.

From Linux performance tuning to building automation with Python/Go, you will collaborate with hardware vendors like NVIDIA and drive scalability for massive distributed training workloads.

Qualifications

  • 5+ years as an SRE/production infra engineer in large-scale environments.
  • Deep Linux, containerized systems, and low-level performance debugging.

Responsibilities

  • Architect for Reliability and Scale across GPU infrastructure.
  • Own multi-cloud GPU clusters (AWS & OCI) with high availability.
  • Drive security and compliance (SOC 2, ISO) measures in startup infra.
  • Perform deep Linux performance tuning at OS/kernel level.
  • Build automation tools in Python, Go, or Bash for infra management.
  • Debug complex hardware/software failures with hardware vendors.

Skills

Linux
Terraform
Airflow
Ray
Kubernetes
Networking
Python
Go
Bash

Tools

AWS
OCI
DCGM/ROCm
InfiniBand/RDMA

Job description

About Luma AI

Luma’s mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries of scale. Our SRE team is the foundation of our research and product velocity, responsible for the thousands of NVIDIA and AMD GPUs across multiple providers that power our work.

Where You Come In

We are looking for a hands-on, first-principles engineer who is fluent in Linux, comfortable operating close to the metal, and capable of architecting systems for the next generation of AI infrastructure.

You will build, maintain, and scale Luma’s infrastructure across on-prem and multi-vendor clouds (AWS & OCI), serving as the bridge between hardware vendors, cloud providers, and our research teams.

What You’ll Do
  • Architect for Reliability & Scale: Participate in critical re-architecture sessions to redesign our systems for higher efficiency and scale. You won't just maintain existing clusters; you will help define how our next-generation infrastructure operates.
  • Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
  • Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices in a fast-moving AI startup environment.
  • Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
  • Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure without relying on heavy operational toil.
  • Debug Complex Hardware/Software Failures: Serve as the final escalation point for the most challenging GPU, networking (InfiniBand/RDMA), and system-level issues, often collaborating directly with hardware vendors like NVIDIA.
Who You Are
  • 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep Linux Mastery: You possess deep, hands‑on expertise in Linux, containerized systems, and debugging low-level system performance.
  • Expert in Technologies: You have working experiencewith Terraform, Airflow, and Ray
  • Cloud Infrastructure Expert: You have strong experience with providers like AWS or OCI.
  • Tenacious Troubleshooter: You thrive on solving complex, low-level problems where hardware and software intersect.
  • Startup DNA: You are energetic and thrive in a less structured, fast‑paced environment.
  • Security-Minded: You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
  • Expert in High-Performance Networking: You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.
What Sets You Apart (Bonus Points)
  • Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.
  • Experience managing large‑scale GPU clusters for AI/ML workloads (training or inference).
  • Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in Data Pipeline and Infrastructure
Compensation

The base pay range for this role is $170,000 – $290,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 360,000
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Head of Global Compute Capacity & Platform Strategy
Head of Global Compute Capacity & Platform Strategy

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000