RL Systems Engineer: Inference & Training at Scale

xAI

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

xAI is seeking an engineer for the RL infrastructure team to help with low precision RL training and inference. You will design and optimize the inference stack for RL workloads, profile performance, and collaborate with the modelling team to implement novel RL techniques.

This role requires building and optimizing large-scale distributed systems and proficiency in Python, C++, or Rust, along with PyTorch, Jax, and CUDA. Strong problem-solving and deep-stack focus are essential for success.

Qualifications

  • Experience in building, debugging, and optimizing efficiency of large-scale distributed systems.
  • Experience in LLM inference.
  • Proficiency in programming languages such as Python, C++ and/or Rust; frameworks such as PyTorch, Jax, CUDA.
  • Willingness to dive deep and solve hardcore problems at all levels of the stack.

Responsibilities

  • Design and optimize the inference stack for RL workloads across sizes.
  • Analyze, profile and address performance bottlenecks in large-scale RL systems.
  • Work closely with the modelling team to implement novel RL techniques and algorithms.

Skills

Distributed systems
LLM inference
Programming languages

Tools

Python
C++
Rust
PyTorch
JAX
CUDA

Job description

xAI is seeking an engineer for the RL infrastructure team to help with low precision RL training and inference. You will design and optimize the inference stack for RL workloads, profile performance, and collaborate with the modelling team to implement novel RL techniques.

This role requires building and optimizing large-scale distributed systems and proficiency in Python, C++, or Rust, along with PyTorch, Jax, and CUDA. Strong problem-solving and deep-stack focus are essential for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Member of Technical Staff - RL Inference
Member of Technical Staff - RL Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - RL Inference
Member of Technical Staff - RL Inference

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000