Member of Technical Staff — Training

RadixArk

Palo Alto (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive benefits
Flexible work arrangements

Job summary

A leading AI infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another systems language, such as C++. This role involves collaborating with researchers and improving the reliability of long-running training jobs. Competitive compensation and equity are offered, alongside comprehensive benefits.

Qualifications

  • 5+ years of experience in ML systems or large-scale training infrastructure.
  • Strong experience with distributed training and performance trade-offs.
  • Proficient in Python and a systems language (C++, Go, Rust).

Responsibilities

  • Design and operate large-scale distributed training systems.
  • Optimize throughput and hardware efficiency.
  • Collaborate with model researchers for experiments.

Skills

Machine Learning systems
Distributed systems
Large-scale training infrastructure
Python
C++

Tools

PyTorch
JAX

Job description

RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models.

You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, and reliability across thousands of GPUs. This role sits at the intersection of ML, systems, and performance engineering.

Your work will directly impact how next-generation AI models are trained and scaled.

This is a deeply technical, high-impact role for engineers who enjoy solving hard systems problems at extreme scale.

Requirements

5+ years of experience in ML systems, distributed systems, or large-scale training infrastructure

Strong experience with large-scale distributed training (data, tensor, and pipeline parallelism)

Deep understanding of GPU/TPU architecture and performance trade-offs

Strong knowledge of PyTorch or JAX distributed training stacks

Experience debugging performance and stability issues in large training jobs

Solid distributed systems fundamentals (networking, consensus, fault tolerance)

Proficiency in Python plus a systems language (C++, Go, or Rust)

Experience operating production ML systems at scale

Strong Plus

Familiarity with DeepSpeed, Megatron-LM, FSDP, or custom training stacks

Experience with RDMA, InfiniBand, or high-speed interconnects

Background in HPC or performance-critical computing

Contributions to ML systems open-source projects

Experience with checkpointing, fault recovery, and elastic training

Experience optimizing training cost efficiency at scale

Responsibilities

Design and operate large-scale distributed training systems

Optimize throughput, scalability, and hardware efficiency

Improve reliability and fault tolerance for long-running training jobs

Develop training frameworks and infrastructure tooling

Collaborate with model researchers to support frontier experiments

Debug and resolve cross-layer performance bottlenecks

Build observability systems for training performance and reliability

Drive capacity planning and cluster utilization strategies

Contribute to long-term training infrastructure architecture

About RadixArk

RadixArk is an infrastructure-first AI company built by engineers who have shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles, our large-scale RL framework.

We build world-class infrastructure for AI training and inference and partner with frontier AI teams and cloud providers.

Our team has coordinated training across 10,000+ GPUs and optimized kernels serving billions of tokens daily.

Join us in building the infrastructure that trains the next generation of AI.

Compensation

We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.

RadixArk is an Equal Opportunity Employer and welcomes candidates from all backgrounds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Cluster / Platform
Member of Technical Staff — Cluster / Platform

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff — Inference
Member of Technical Staff — Inference

RadixArk • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Member of Technical Staff — Inference-Kernel, Compiler & Communication
Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Member of Technical Staff — Heterogenous Hardware
Member of Technical Staff — Heterogenous Hardware

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Flexible work arrangements
+1
Business Development
Business Development

RadixArk • Palo Alto (CA)

On-site
USD 150,000 - 200,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
AI Infra Resident (1-Year Program)
AI Infra Resident (1-Year Program)

RadixArk • Palo Alto (CA)

On-site
USD 70,000 - 90,000
Health benefits
Potential for full-time position with equity
Member of Technical Staff — Inference-Multimodal & Diffusion
Member of Technical Staff — Inference-Multimodal & Diffusion

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Talent Operations
Talent Operations

RadixArk • Palo Alto (CA)

Hybrid
USD 118,000 - 145,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000