Senior Software Engineer, GPU Cluster Infrastructure

Far Ai

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

FAR.AI is a nonprofit AI safety research institute. Foundations, its infrastructure and engineering team, builds the compute platform, tools, and automation enabling researchers.

The role spans scheduling, storage, monitoring, and security for large GPU clusters, with emphasis on pre-training and post-training infrastructure and multi-tenancy to keep experiments fast and safe. Work involves collaborating with researchers and engineers, scaling infrastructure as the organization grows, and

Qualifications

  • Experience operating large-scale GPU clusters and pre-/post-training workflows.
  • Strong background in distributed storage and networking for HPC.
  • Proficiency with infrastructure as code and automation tools.

Responsibilities

  • Operate the Kubernetes GPU fleet day to day, including upgrades and rollout.
  • Own batch scheduling, multi-tenancy, queues, quotas, and priorities.
  • Design and manage storage for datasets, checkpoints, and backups.
  • Maintain fault-tolerant multi-node training and cluster health.
  • Harden platform security, identity, network policy, and sandboxing.
  • Onboard new capacity, test providers, and integrate into IaC pipelines.
  • Collaborate directly with researchers and engineers across teams.

Skills

Kubernetes
GPU clusters
Batch scheduling
Security posture
Infrastructure as code

Education

BS in Computer Science or related field

Tools

NVIDIA NCCL
Terraform
Kubernetes tooling

Job description

About Us

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.

We're structured to support that work from early research through real-world adoption:

Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly.

A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.

Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.

Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.

Since our founding in July 2022, we've grown to 50+ staff, published 40+ academic papers, and convened leading AI safety events. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026, and ICLR, and features in the Financial Times, Nature News, Wired Magazine and MIT Technology Review. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office and publish the AI Security Leaderboard based on our red-team expertise. We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers.

About the Team

Foundations is FAR.AI's infrastructure and engineering team. Our remit is broad: we run the compute platform, build the tools and frameworks researchers work in, automate research workflows, and help teams scale experiments well past what they'd manage alone. Our job is to accelerate the research. We do so by working directly with researchers through embedded engagements and day-to-day consulting, and building systems that can scale with the organization as it grows.

Foundations is growing quickly, and our infrastructure portfolio is growing fastest. We run FAR.AI's research on a mix of bare-metal and managed Kubernetes GPU clusters from multiple providers. We rent the hardware and operate the platform ourselves. The fleet has grown from dozens to hundreds of GPUs this year and it's continuing to grow quickly: we're adding providers, taking on users beyond our own researchers, and moving experiments onto frontier open-weight models. A large amount of research is now being done by AI agents working directly on the cluster, which is driving updates to our platform infrastructure and security.

Running it well now takes dedicated specialists, so we're standing up an infrastructure sub-team that owns the cluster fleet: adding capacity, designing and managing the networking and storage under it, infrastructure as code, and the security posture, plus some of the platform layer above it. It works directly with research teams as their needs change.

About the Role

You'd work across the whole infrastructure stack, from scheduling to storage to monitoring to security, and bring real depth in at least one part of it. We're particularly interested in experience with large-scale pre-training and post-training infrastructure and the network fabric under it, cluster security and sandboxing, distributed storage systems, and batch scheduling for large GPU clusters. Expertise in an adjacent area is also a good fit.

In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant.

We're also hiring a Tech Lead Manager, GPU Cluster Infrastructure for this team. If leading a small team while staying hands-on sounds like you, take a look there instead.

What you'll do
  • Operate the Kubernetes GPU fleet day to day. You handle node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.

  • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams.

  • Design and run the storage under the fleet, from high-performance shared filesystems for datasets and checkpoints to object storage tiers, quotas, and backups.

  • Keep multi-node training runs fault-tolerant. You own node health and automated draining, debug NCCL and fabric problems, track down stragglers and flaky GPUs, and build the checkpoint and restart patterns.

  • Harden the platform, covering identity and access, network policy, secrets, workload isolation, and sandboxing for the AI agents that run on the cluster.

  • Bring new capacity online. You acceptance-test providers on fabric, NCCL, and storage throughput, hold them to their SLAs, and integrate new clusters into the platform with infrastructure as code.

  • Work directly

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead Manager, GPU Cluster Infrastructure
Tech Lead Manager, GPU Cluster Infrastructure

Far Ai • United States

On-site
USD 180,000 - 250,000
Senior Software Engineer, GPU Cluster Infrastructure
Senior Software Engineer, GPU Cluster Infrastructure

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Tech Lead Manager, GPU Cluster Infrastructure
Tech Lead Manager, GPU Cluster Infrastructure

AISafety • Berkeley (CA)

On-site
USD 180,000 - 240,000
Health Insurance
401(k) plan
PTO - 25 days per year
+4
GPU Cluster Infra Tech Lead & Team Builder
GPU Cluster Infra Tech Lead & Team Builder

AISafety • Berkeley (CA)

Hybrid
USD 180,000 - 240,000
Health Insurance
401(k) plan
PTO - 25 days per year
+4
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior GPU Infrastructure Engineer for AI Research
Senior GPU Infrastructure Engineer for AI Research

Far Ai • United States

Remote
USD 180,000 - 240,000
GPU Cluster Infra Lead — Tech Strategy & Team Growth
GPU Cluster Infra Lead — Tech Strategy & Team Growth

Far Ai • United States

Remote
USD 180,000 - 250,000
Member of Technical Staff, AI Compute & Data Infrastructure
Member of Technical Staff, AI Compute & Data Infrastructure

Vinci • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Technical Program Manager, Research
Technical Program Manager, Research

FAR.AI • Berkeley (CA)

Hybrid
USD 120,000 - 190,000
Health Insurance
401(k) match
PTO 25 days/yr
+3