GPU Cluster Infra Tech Lead & Team Builder

AISafety

Berkeley (CA)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) plan
PTO - 25 days per year
Paid Bereavement, Family, Medical and
Pregnancy Disability Leave
WFH stipend & Equipment
Catered meals (Berkeley Office)

Job summary

FAR.AI in Berkeley, CA, is seeking an infrastructure leader to own the platform's technical direction for GPU-heavy research workloads. You will partner with research teams, hire and mentor senior engineers, and shape the roadmap for a scalable, fault-tolerant compute cluster.

This role emphasizes hands-on engineering alongside leadership, with opportunities to influence security posture, on-call processes, and cross-provider orchestration as FAR.AI scales its frontier AI safety work.

Qualifications

  • You've led engineers as a manager, tech lead, or project lead, setting technical direction, scoping work, and giving feedback.
  • You have 5+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, and you've owned systems from design through operation.
  • You've run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health.
  • You've owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents.
  • You're a strong programmer in at least one language that infrastructure is commonly written in, such as Python, Go, Rust, or C++, and your automation and services are maintained as shared code.
  • You write clearly for engineers, researchers, and providers, whether it's a roadmap, a design doc, an incident summary, or an escalation.

Responsibilities

  • Set the platform's technical direction and own its roadmap. You decide which systems we run, how we schedule and store across providers, and what we measure.
  • Stay in the work. You own architecture and the scheduling and storage design, and you debug the failures that cross layers, such as node health, GPU and fabric faults, and multi-node job hangs.
  • Hire and grow a small team of senior engineers. You set priorities and ownership, scope projects, and give regular feedback and coaching.
  • Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing.
  • Decide how we operate, from the on-call rotation and incident response to postmortems, fault tolerance, and observability, and take part in it alongside the team.
  • Be the escalation point for research teams and for providers, and turn recurring problems into platform fixes.

Skills

Leadership experience
Systems/infrastructure engineering
Kubernetes GPU workloads
Terraform/Ansible
Programming: Python/Go/Rust/C++
Technical writing

Tools

Kubernetes
Slurm/Kueue/Volcano
Terraform
Ansible
Helm
ArgoCD
Prometheus

Job description

FAR.AI in Berkeley, CA, is seeking an infrastructure leader to own the platform's technical direction for GPU-heavy research workloads. You will partner with research teams, hire and mentor senior engineers, and shape the roadmap for a scalable, fault-tolerant compute cluster.

This role emphasizes hands-on engineering alongside leadership, with opportunities to influence security posture, on-call processes, and cross-provider orchestration as FAR.AI scales its frontier AI safety work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Engineering Manager, GPU Infrastructure & Platforms
Engineering Manager, GPU Infrastructure & Platforms

cohere • United States

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K matching
+5
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000