Staff ML Systems Engineer — Distributed Training at Scale
RadixArk
Palo Alto (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive benefits
Flexible work arrangements
Job summary
A leading AI infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another systems language, such as C++. This role involves collaborating with researchers and improving the reliability of long-running training jobs. Competitive compensation and equity are offered, alongside comprehensive benefits.
Qualifications
5+ years of experience in ML systems or large-scale training infrastructure.
Strong experience with distributed training and performance trade-offs.
Proficient in Python and a systems language (C++, Go, Rust).
Responsibilities
Design and operate large-scale distributed training systems.
Optimize throughput and hardware efficiency.
Collaborate with model researchers for experiments.
Skills
Machine Learning systems
Distributed systems
Large-scale training infrastructure
Python
C++
Tools
PyTorch
JAX
Job description
A leading AI infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another systems language, such as C++. This role involves collaborating with researchers and improving the reliability of long-running training jobs. Competitive compensation and equity are offered, alongside comprehensive benefits.