RL Infrastructure Engineer — Scalable Training & Performance

xAI

Palo Alto (CA)

On-site

USD 170,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Life and AD&D insurance
Fertility benefits
Flexible vacation
Visa sponsorship
401(k) plan

Job summary

xAI in Palo Alto is seeking an experienced engineer to design and implement the RL training framework and the systems backing all RL workloads, from ablations to production runs. You will profile, debug, and optimize end-to-end training performance, and improve scalability and observability of the RL stack.

The role requires proficiency in Python and C++, familiarity with Jax or Rust, and experience with large-scale RL/NLP training infrastructures and RL numerics.

Qualifications

  • Experience building, debugging, and optimizing efficiency of large-scale distributed systems.
  • Comfortable diving into unfamiliar areas and solving problems at all levels of the stack.
  • Proficiency in Python, Jax, Rust, and/or C++.
  • Experience with large scale LLM training infrastructure.
  • Strong knowledge of reinforcement learning techniques.
  • Experience with RL numerics.

Responsibilities

  • Design and implement the RL training framework and the systems backing RL workloads, from small-scale ablations to production training runs.
  • Profile, debug, and optimize end-to-end training performance.
  • Improve scalability and observability of the RL stack.
  • Collaborate with the RL infrastructure team to ship reliable, scalable systems.

Skills

Python
Jax
Rust
C++
Reinforcement learning techniques
RL numerics

Job description

xAI in Palo Alto is seeking an experienced engineer to design and implement the RL training framework and the systems backing all RL workloads, from ablations to production runs. You will profile, debug, and optimize end-to-end training performance, and improve scalability and observability of the RL stack.

The role requires proficiency in Python and C++, familiarity with Jax or Rust, and experience with large-scale RL/NLP training infrastructures and RL numerics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - RL Training Infrastructure
Staff Engineer - RL Training Infrastructure

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Large-Scale RL Infrastructure Engineer
Large-Scale RL Infrastructure Engineer

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff RL Systems Engineer — Equity
Staff RL Systems Engineer — Equity

Neura Market • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Member of Technical Staff - RL Training Framework
Member of Technical Staff - RL Training Framework

Xai • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
RL Systems Engineer: Scale Training Pipelines & GPUs
RL Systems Engineer: Scale Training Pipelines & GPUs

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000