Senior Cluster Infra Architect for Frontier AI — Equity

RadixArk

Palo Alto (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

RadixArk is seeking a Member of Technical Staff to architect and scale the core compute platform powering frontier-level AI training and inference. You will design and operate GPU/TPU clusters, build scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.

This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization, with impact on how efficiently frontier AI

Qualifications

  • 5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms.
  • Strong background in distributed systems design and systems architecture.
  • Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers).
  • Hands-on experience with GPU/TPU infrastructure in production environments.
  • Strong Linux systems and networking fundamentals.
  • Proficiency in Go, Rust, C++, or Python for production systems.
  • Experience debugging complex multi-layer issues across hardware, OS, networking, and distributed services.
  • Proven ability to design reliable, scalable systems in production.

Responsibilities

  • Architect and scale large AI compute clusters for training and inference.
  • Design cluster management, scheduling, and resource allocation systems.
  • Optimize performance, utilization, and reliability of GPU/TPU clusters.
  • Improve fault tolerance and system resilience at scale.
  • Drive observability, monitoring, and performance profiling for cluster infrastructure.
  • Collaborate with ML and systems engineers to support frontier AI workloads.
  • Lead capacity planning and infrastructure scaling strategies.
  • Build internal platforms and tooling to improve developer productivity.
  • Document architecture, operational practices, and reliability strategies.
  • Contribute to long-term platform vision and technical direction.

Skills

Distributed systems
Cluster management
Kubernetes
Slurm
Ray
Go
Rust
C++
Python
Linux networking
GPU/TPU infrastructure
Production systems
Debugging complex systems

Tools

Kubernetes
Slurm
Ray
Custom schedulers

Job description

RadixArk is seeking a Member of Technical Staff to architect and scale the core compute platform powering frontier-level AI training and inference. You will design and operate GPU/TPU clusters, build scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.

This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization, with impact on how efficiently frontier AI

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Cluster / Platform
Member of Technical Staff — Cluster / Platform

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Senior GPU Infrastructure Architect for Frontier AI
Senior GPU Infrastructure Architect for Frontier AI

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff — Training
Member of Technical Staff — Training

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Staff Engineer, Frontier AI Infrastructure
Staff Engineer, Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range: $150-300k
Flexible work arrangement (SF office +
Full visa sponsorship and relocation
+1
Principal Software Engineer: Frontier AI Infra & Platform
Principal Software Engineer: Frontier AI Infra & Platform

NVIDIA • Santa Clara (CA)

On-site
USD 248,000 - 391,000
Equity
Benefits
Member of Technical Staff — Inference
Member of Technical Staff — Inference

RadixArk • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Member of Technical Staff — Inference-Kernel, Compiler & Communication
Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Staff Storage Infrastructure Engineer – Frontier AI
Staff Storage Infrastructure Engineer – Frontier AI

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Remote RL Infrastructure Engineer - Frontier AI Stack
Remote RL Infrastructure Engineer - Frontier AI Stack

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites