Member of Technical Staff - RL Training Framework

SpaceXAI

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SpaceXAI seeks a Member of Technical Staff to design and implement the RL training framework and supporting systems. You will work on end-to-end RL workloads, from ablations to production training runs, profiling performance, and improving scalability and observability across the stack.

The role requires deep experience with distributed systems and proficiency in Python, Jax, Rust, and/or C++, with preferred familiarity in RL algorithms and large-scale training infrastructure.

Qualifications

  • Experience building, debugging, and optimizing efficiency of large-scale distributed systems.
  • Comfortable diving into unfamiliar areas and solving problems at all levels of the stack.
  • Proficiency in Python, Jax, Rust, and/or C++.

Responsibilities

  • Design and implement the systems backing all RL workloads at SpaceXAI, from small scale ablations to production training runs.
  • Profile, debug, and optimize end-to-end training performance
  • Improve scalability and observability of the RL stack

Skills

Distributed systems
Performance optimization
Problem solving
Communication

Tools

Python
Jax
Rust
C++

Job description

Member of Technical Staff - RL Training Framework

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands‑on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE

The RL infrastructure team is looking for an engineer to help develop our RL training framework.

RESPONSIBILITIES
  • Design and implement the systems backing all RL workloads at SpaceXAI, from small scale ablations to production training runs.
  • Profile, debug, and optimize end-to-end training performance
  • Improve scalability and observability of the RL stack
BASIC QUALIFICATIONS
  • Experience building, debugging, and optimizing efficiency of large-scale distributed systems
  • Comfortable diving into unfamiliar areas and solving problems at all levels of the stack
  • Proficiency in Python, Jax, Rust, and/or C++
PREFERRED SKILLS AND EXPERIENCE
  • Experience with large scale LLM training infrastructure
  • Strong knowledge of reinforcement learning techniques
  • Experience with RL numerics
COMPENSATION AND BENEFITS

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - RL Training Framework
Member of Technical Staff - RL Training Framework

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Large-Scale RL Infrastructure Engineer
Large-Scale RL Infrastructure Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Operations Engineer - Human Engineer
Operations Engineer - Human Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Equity
Medical Insurance
Vision & Dental
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Software Engineer - Platform Infrastructure (Rust, C++)
Software Engineer - Platform Infrastructure (Rust, C++)

SpaceXAI • Bellevue (WA)

On-site
USD 180,000 - 440,000
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Machine Learning Engineer - Recommendation Systems
Machine Learning Engineer - Recommendation Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and/or
Dental coverage
+4
Sr. Software Engineer (Vehicle Engineering)
Sr. Software Engineer (Vehicle Engineering)

jobs.frontdoordefense.com - Jobboard • Hawthorne (CA)

On-site
USD 160,000 - 225,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Member of Technical Staff - Multimodal Understanding
Member of Technical Staff - Multimodal Understanding

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical insurance
Vision insurance
+4