Senior GPU Cluster Infra Engineer | Remote

AISafety

Berkeley (CA)

Hybrid

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) match
PTO 25 days per year
Bereavement leave
WFH stipend
Catered meals (Berkeley Office)

Job summary

FAR.AI seeks an experienced infrastructure engineer to manage and scale its GPU cluster infrastructure, spanning scheduling, storage, and security. You will collaborate with researchers and other engineers to ensure high availability and efficient workload execution on large-scale GPU clusters.

The role emphasizes hands-on work with Kubernetes, batch systems, and secure multi-tenant environments, contributing to frontier AI research while maintaining robust production operations.

Qualifications

  • 3+ years in systems or infrastructure engineering on production Linux
  • Experience running GPU workloads with a batch layer
  • Own infrastructure as code and observability for a production fleet
  • Proficient programming in Python, Go, Rust, or C++
  • Able to write clearly for engineers, researchers, and providers

Responsibilities

  • Operate the Kubernetes GPU fleet day to day, including node lifecycle, upgrades, and capacity planning
  • Own batch scheduling, quotas, priorities, preemption, and fair share across teams
  • Design and run storage under the fleet with high-performance shared filesystems and object storage
  • Maintain fault tolerance for multi-node training, diagnose NCCL/fabric issues, and implement checkpoint/restart patterns
  • Harden the platform with identity, network policy, secrets, and sandboxing for AI agents

Skills

Kubernetes admin
GPU workloads
Infrastructure as code
Programming (Python/Go/C++)
System scalability

Tools

Terraform
Ansible
Helm
ArgoCD
Prometheus

Job description

FAR.AI seeks an experienced infrastructure engineer to manage and scale its GPU cluster infrastructure, spanning scheduling, storage, and security. You will collaborate with researchers and other engineers to ensure high availability and efficient workload execution on large-scale GPU clusters.

The role emphasizes hands-on work with Kubernetes, batch systems, and secure multi-tenant environments, contributing to frontier AI research while maintaining robust production operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Infra Tech Lead & Team Builder
GPU Cluster Infra Tech Lead & Team Builder

AISafety • Berkeley (CA)

Hybrid
USD 180,000 - 240,000
Health Insurance
401(k) plan
PTO - 25 days per year
+4
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Compute Cluster Engineer
Senior GPU Compute Cluster Engineer

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110,000 - 150,000
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000