Tech Lead Manager, GPU Cluster Infrastructure

AISafety

Berkeley (CA)

Hybrid

USD 180,000 - 240,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) plan
PTO - 25 days per year
Paid Bereavement, Family, Medical and
Pregnancy Disability Leave
WFH stipend & Equipment
Catered meals (Berkeley Office)

Job summary

FAR.AI in Berkeley, CA, is seeking an infrastructure leader to own the platform's technical direction for GPU-heavy research workloads. You will partner with research teams, hire and mentor senior engineers, and shape the roadmap for a scalable, fault-tolerant compute cluster.

This role emphasizes hands-on engineering alongside leadership, with opportunities to influence security posture, on-call processes, and cross-provider orchestration as FAR.AI scales its frontier AI safety work.

Qualifications

  • You've led engineers as a manager, tech lead, or project lead, setting technical direction, scoping work, and giving feedback.
  • You have 5+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, and you've owned systems from design through operation.
  • You've run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health.
  • You've owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents.
  • You're a strong programmer in at least one language that infrastructure is commonly written in, such as Python, Go, Rust, or C++, and your automation and services are maintained as shared code.
  • You write clearly for engineers, researchers, and providers, whether it's a roadmap, a design doc, an incident summary, or an escalation.

Responsibilities

  • Set the platform's technical direction and own its roadmap. You decide which systems we run, how we schedule and store across providers, and what we measure.
  • Stay in the work. You own architecture and the scheduling and storage design, and you debug the failures that cross layers, such as node health, GPU and fabric faults, and multi-node job hangs.
  • Hire and grow a small team of senior engineers. You set priorities and ownership, scope projects, and give regular feedback and coaching.
  • Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing.
  • Decide how we operate, from the on-call rotation and incident response to postmortems, fault tolerance, and observability, and take part in it alongside the team.
  • Be the escalation point for research teams and for providers, and turn recurring problems into platform fixes.

Skills

Leadership experience
Systems/infrastructure engineering
Kubernetes GPU workloads
Terraform/Ansible
Programming: Python/Go/Rust/C++
Technical writing

Tools

Kubernetes
Slurm/Kueue/Volcano
Terraform
Ansible
Helm
ArgoCD
Prometheus

Job description

About Us

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.


We’re structured to support that work from early research through real-world adoption:



  • Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly.

  • A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.

  • Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.

  • Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.


Since our founding in July 2022, we've grown to 50+ staff, published 40+ academic papers, and convened leading AI safety events. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026, and ICLR, and features in the Financial Times, Nature News, Wired Magazine and MIT Technology Review. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office and publish the AI Security Leaderboard based on our red‑teaming expertise. We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs, an AI safety-focused co‑working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers.


About the Team

Foundations is FAR.AI's infrastructure and engineering team. Our remit is broad: we run the compute platform, build the tools and frameworks researchers work in, automate research workflows, and help teams scale experiments well past what they'd manage alone. Our job is to accelerate the research. We do so by working directly with researchers through embedded engagements and day‑to‑day consulting, and building systems that can scale with the organization as it grows.


Foundations is growing quickly, and our infrastructure portfolio is growing fastest. We run FAR.AI's research on a mix of bare‑metal and managed Kubernetes GPU clusters from multiple providers. We rent the hardware and operate the platform ourselves. The fleet has grown from dozens to hundreds of GPUs this year and it's continuing to grow quickly: we're adding providers, taking on users beyond our own researchers, and moving experiments onto frontier open‑weight models. A large amount of research is now being done by AI agents working directly on the cluster, which is driving updates to our platform infrastructure and security.


Running it well now takes dedicated specialists, so we're standing up an infrastructure sub‑team that owns the cluster fleet: adding capacity, designing and managing the networking and storage under it, infrastructure as code, and the security posture, plus some of the platform layer above it. It works directly with research teams as their needs change.


About the Role

You'd be the infrastructure sub‑team's lead and one of its engineers, with at least half your time on technical work. You'd own the platform's technical direction and roadmap, deciding what we build and how we run it in collaboration with the research teams and the rest of Foundations, and you'd hire and grow the team.


In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large‑scale experiments performant and fault‑tolerant.


We're also hiring a Software Engineer, GPU Cluster Infrastructure for this team. If you want the hands‑on work without the management responsibilities, take a look there instead.


What you'll do


  • Set the platform's technical direction and own its roadmap. You decide which systems we run, how we schedule and store across providers, and what we measure, from utilization and queue wait to failure rates.

  • Stay in the work. You own architecture and the scheduling and storage design, and you debug the failures that cross layers, such as node health, GPU and fabric faults, and multi‑node job hangs.

  • Hire and grow a small team of senior engineers. You set priorities and ownership, scope projects, and give regular feedback and coaching.

  • Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing.

  • Decide how we operate, from the on‑call rotation and incident response to postmortems, fault tolerance, and observability, and take part in it alongside the team.

  • Be the escalation point for research teams and for providers, and turn recurring problems into platform fixes.


Requirements


  • You've led engineers as a manager, tech lead, or project lead, setting technical direction, scoping work, and giving feedback. You want to manage people directly; prior direct reports aren't required.

  • You have 5+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large‑scale batch platforms, and you've owned systems from design through operation. You have depth in at least one of scheduling, storage, networking, security, or GPU systems, and enough breadth to review designs in the rest.

  • You've run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health.

  • You've owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents.

  • You're a strong programmer in at least one language that infrastructure is commonly written in, such as Python, Go, Rust, or C++, and your automation and services are maintained as shared code.

  • You write clearly for engineers, researchers, and providers, whether it's a roadmap, a design doc, an incident summary, or an escalation.


If you meet most of this and not all of it, we encourage you to apply anyway.


Additional skills we're excited about

Real depth in one or more of these makes you a compelling candidate:



  • Distributed training infrastructure: multi‑node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology‑aware placement.

  • Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O.

  • Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure.

  • Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair‑share and quota, and the utilization, fairness, and latency tradeoffs between them.

  • Multi‑provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality.

  • Greenfield team‑building: you've hired senior engineers and stood up on‑call, incident, and review practices for a new team before.


Benefits*


  • Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date

  • Retirement - 401(k) plan with up to 2% match

  • PTO - 25 days Paid Time Off per year, accrued weekly and up to 10 days of paid sick leave per year

  • Paid Leave - Paid Bereavement, Family, Medical and Pregnancy Disability Leave

  • WFH Stipend & Equipment- Work computer and stipend provided for eligible employees

  • Catered Meals (Berkeley Office Only) - Catered lunches and dinners on workdays at our office


* (Available only to full‑time employees located in the US)


Logistics

If based in the USA or Singapore, you will be an employee of FAR.AI (501(c)(3) research non‑profit / non‑profit CLG). Outside the USA or Singapore, you will be employed via an EOR organisation on behalf of FAR.AI or as a contractor.



  • Location: Both remote and in‑person (Berkeley, CA or Singapore) are possible. We sponsor visas for in‑person employees, and can hire remotely in most countries. For this role we prefer candidates whose working hours overlap with Berkeley.

  • Hours: Full‑time (40 hours/week).

  • On‑call: We don't run a formal on‑call rotation yet. The team is spread across time zones and covers incidents during working hours. As the experiments we run get larger we expect to introduce one, and this role would take part in it.

  • Application process: A one hour programming assessment, a short screening call, a one hour code review interview, a 90 minute systems design and incident response interview, a one hour management and leadership interview, and a paid work trial lasting up to 1 week. If you are not available for a work trial we may be able to find alternative ways of testing your fit.


If you have any questions about the role, feel free to contact us at talent@far.ai.


Please don't email us to share your resume (it won't have any impact on our decision). Thank you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, GPU Cluster Infrastructure
Senior Software Engineer, GPU Cluster Infrastructure

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Tech Lead Manager, GPU Cluster Infrastructure
Tech Lead Manager, GPU Cluster Infrastructure

FAR.AI • United States

Hybrid
USD 225,000 - 325,000
Health Insurance
401(k) Match
Paid Time Off
+4
IT and Security Manager
IT and Security Manager

AISafety • Berkeley (CA)

On-site
USD 140,000 - 220,000
Health Insurance
Retirement match
Paid time off
+3
IT & Security Manager
IT & Security Manager

FAR.AI • United States

On-site
USD 185,000 - 250,000
Health Insurance
401(k) match
25 days PTO
+3
Research Scientist
Research Scientist

Aisafety • Berkeley (CA)

Hybrid
USD 120,000 - 190,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Reimbursement for work-related travel
Senior Research Engineer
Senior Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 150,000 - 250,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Work-related travel expenses covered
Technical Program Manager, Research
Technical Program Manager, Research

Aisafety • Berkeley (CA)

Hybrid
USD 125,000 - 190,000
Technical Program Manager, Research
Technical Program Manager, Research

FAR.AI • Berkeley (CA)

Hybrid
USD 125,000 - 190,000
Visa sponsorship for US work
Remote/U.S. location flexibility
Competitive compensation USD 125,000–$
Research Engineer
Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 100,000 - 190,000
Catered lunch and dinner
Potential for remote work
Sponsorship for work-related travel
Research Lead
Research Lead

Aisafety • Berkeley (CA)

Hybrid
USD 170,000 - 270,000
Catered lunch and dinner
Visa sponsorship