Member of Technical Staff, Cluster Infrastructure

Goaly

Menlo Park, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Meals and office benefits
Visa sponsorship

Job summary

Goaly is seeking an experienced infrastructure engineer to own the lifecycle of accelerator clusters across cloud and datacenter environments. You will design and operate control-plane services, provisioning, upgrades, and decommissioning while partnering with multiple teams to turn heterogeneous compute into a dependable platform.

You will work on a hands-on, hybrid role spanning networking, storage, security, and automation.

Qualifications

  • Experience in distributed systems, cloud platforms, and reliable infra services.
  • Experience with Kubernetes, IaC, and at least one major cloud provider.
  • Strong programming ability in Python, Go, or Rust for automation and tooling.
  • Hands-on Linux, containers, and cluster schedulers; practical networking and observability.

Responsibilities

  • Design, build, and operate control-plane services and infrastructure-as-code for provisioning, upgrades, and recovery.
  • Bring accelerator capacity online across cloud and datacenter environments.
  • Improve fault tolerance and scalability with automated remediation and health signals.
  • Ensure security by default with identity controls, network policy, and trusted software supply chains.
  • Establish observability for cluster readiness, capacity, and recovery times.
  • Lead incident response and blameless postmortems for cluster failures.
  • Collaborate with Training, RL Systems, Post-Training, and Inference teams to shape compute and capacity roadmap.

Skills

Distributed systems
Kubernetes
Cloud platforms
Python/Go/Rust
Linux
IaC
Observability
Networking

Tools

Terraform
Kubernetes
Slurm

Job description

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.

About the role

You will own systems across the lifecycle of our accelerator clusters—from bringing capacity online and upgrading fleets to detecting failures, recovering safely, and retiring capacity. Your work will determine how quickly researchers can start experiments, how efficiently expensive hardware is used, and how reliably long-running workloads complete.

This is a hands-on infrastructure role spanning cloud and datacenter environments, cluster control planes, networking, storage, security, observability, and automation. You will partner with hardware and cloud providers and our Training, RL Systems, Post-Training, Inference, and Security teams to turn heterogeneous compute into a dependable platform. Depending on experience, you may lead multi-quarter initiatives and help set technical direction.

What you'll do
  • Design, build, and operate control-plane services and infrastructure-as-code for provisioning, configuration, validation, upgrades, expansion, draining, recovery, and decommissioning; make every change repeatable, auditable, and safe to roll back.

  • Bring new accelerator capacity online on schedule by coordinating dependencies across cloud providers, datacenter and hardware partners, networking, storage, security, and internal compute consumers.

  • Build high-bandwidth, topology-aware connectivity within and across clusters; diagnose performance and reliability issues spanning hosts, switches, routing, transport, collective communication, and workload placement.

  • Make clusters secure by default through identity and access controls, network policy, workload isolation, host and container hardening, secrets management, and trusted software and image supply chains.

  • Improve fleet scalability, consistency, and fault tolerance by defining health signals, automating remediation, reducing configuration drift, and designing for partial failure.

  • Establish service-level objectives and observability for cluster readiness, provisioning time, usable capacity, job-start latency, infrastructure-caused failures, utilization, and recovery time.

  • Lead incident response and blameless postmortems for cluster failures; turn recurring operational pain into automation, safer defaults, and simpler system boundaries.

  • Work directly with Training, RL Systems, Post-Training, and Inference engineers to debug cross-layer failures and shape a long-term compute, data, networking, and capacity roadmap.

You may be a good fit if you have
  • Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure).

  • Strong programming ability in Python, Go, Rust, or another language suited to reliable infrastructure services and automation.

  • Hands-on experience with Linux, containers, Kubernetes or another cluster scheduler, infrastructure-as-code, and at least one major cloud platform or substantial bare-metal environment.

  • A practical understanding of networking, storage, identity, observability, and reliability, with the ability to trace a failure across multiple layers of a complex system.

  • Experience designing systems for safe rollout, fault isolation, idempotency, capacity growth, and recovery from partial or large-scale failures.

  • High ownership and clear communication, including comfort coordinating multi-team projects and participating in a healthy on-call rotation.

Strong pluses
  • Experience operating large GPU or accelerator fleets for distributed model training, inference, scientific computing, or another communication-intensive workload.

  • Depth in Kubernetes internals, custom controllers or operators, device plugins, cluster autoscaling, scheduler extensions, Slurm, or comparable orchestration systems.

  • Experience with high-performance networking such as RDMA, InfiniBand, RoCE, BGP, cloud interconnects, multi-NIC hosts, CNI or eBPF networking, or topology-aware placement.

  • Experience with Terraform, workflow orchestration, and automated qualification of hosts, drivers, firmware, networks, and new hardware.

  • Knowledge of GPU systems, NCCL, NVLink or NVSwitch, and the failure modes of large distributed jobs.

How we work
  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.
  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.
  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.
  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.
  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.
  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.
Location, visa sponsorship & benefits
  • Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.
  • Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.
  • Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.
A note on qualifications.

We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.

Equal opportunity

We are an equal opportunity employer and consider qualified applicants without regard to any characteristic protected by applicable law. Reasonable accommodations are available throughout the hiring process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Founding Senior AI Infrastructure Engineer
Founding Senior AI Infrastructure Engineer

Goaly • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA
Staff Software Engineer, Node Infra San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000