Senior GPU Infrastructure Engineer for AI Research

Far Ai

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

FAR.AI is a nonprofit AI safety research institute. Foundations, its infrastructure and engineering team, builds the compute platform, tools, and automation enabling researchers.

The role spans scheduling, storage, monitoring, and security for large GPU clusters, with emphasis on pre-training and post-training infrastructure and multi-tenancy to keep experiments fast and safe. Work involves collaborating with researchers and engineers, scaling infrastructure as the organization grows, and

Qualifications

  • Experience operating large-scale GPU clusters and pre-/post-training workflows.
  • Strong background in distributed storage and networking for HPC.
  • Proficiency with infrastructure as code and automation tools.

Responsibilities

  • Operate the Kubernetes GPU fleet day to day, including upgrades and rollout.
  • Own batch scheduling, multi-tenancy, queues, quotas, and priorities.
  • Design and manage storage for datasets, checkpoints, and backups.
  • Maintain fault-tolerant multi-node training and cluster health.
  • Harden platform security, identity, network policy, and sandboxing.
  • Onboard new capacity, test providers, and integrate into IaC pipelines.
  • Collaborate directly with researchers and engineers across teams.

Skills

Kubernetes
GPU clusters
Batch scheduling
Security posture
Infrastructure as code

Education

BS in Computer Science or related field

Tools

NVIDIA NCCL
Terraform
Kubernetes tooling

Job description

FAR.AI is a nonprofit AI safety research institute. Foundations, its infrastructure and engineering team, builds the compute platform, tools, and automation enabling researchers.

The role spans scheduling, storage, monitoring, and security for large GPU clusters, with emphasis on pre-training and post-training infrastructure and multi-tenancy to keep experiments fast and safe. Work involves collaborating with researchers and engineers, scaling infrastructure as the organization grows, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior Software Engineer, GPU Cluster Infrastructure
Senior Software Engineer, GPU Cluster Infrastructure

Far Ai • United States

On-site
USD 180,000 - 240,000
GPU Cluster Infra Tech Lead & Team Builder
GPU Cluster Infra Tech Lead & Team Builder

AISafety • Berkeley (CA)

Hybrid
USD 180,000 - 240,000
Health Insurance
401(k) plan
PTO - 25 days per year
+4
Tech Lead Manager, GPU Cluster Infrastructure
Tech Lead Manager, GPU Cluster Infrastructure

Far Ai • United States

On-site
USD 180,000 - 250,000
GPU Cluster Infra Lead — Tech Strategy & Team Growth
GPU Cluster Infra Lead — Tech Strategy & Team Growth

Far Ai • United States

Remote
USD 180,000 - 250,000
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
Senior Security Engineer: AI Infrastructure & GPU Compute
Senior Security Engineer: AI Infrastructure & GPU Compute

The Consensus • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Health plan
Dental & Vision
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000