AI Training Systems Engineer: Distributed & RL

B Capital

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision insurance
Fully paid parental leave
Daily lunches and dinners provided
Relocation support and paid time off

Job summary

B Capital in San Francisco is looking for an engineering professional to architect and optimize core training infrastructure for their AI models. You will work on distributed systems and large-scale data pipelines, focusing on performance and numerical stability. Successful candidates will have strong software engineering skills and experience in either distributed training or data infrastructure. The role offers top-tier compensation and comprehensive health and wellness benefits.

Qualifications

  • Strong software engineering skills with knowledge of machine learning.
  • Experience in distributed systems and performance optimization.
  • Ability to implement research papers into scalable systems.

Responsibilities

  • Designing and optimizing large-scale training loops and data pipelines.
  • Implementing state-of-the-art techniques ensuring numerical stability.
  • Building internal tooling for launching and monitoring experiments.

Skills

Distributed Training & Inference
Data Infrastructure

Tools

PyTorch
JAX
Kubernetes
Ray

Job description

B Capital in San Francisco is looking for an engineering professional to architect and optimize core training infrastructure for their AI models. You will work on distributed systems and large-scale data pipelines, focusing on performance and numerical stability. Successful candidates will have strong software engineering skills and experience in either distributed training or data infrastructure. The role offers top-tier compensation and comprehensive health and wellness benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Software Engineer — Scalable RL & Distributed Training
Research Software Engineer — Scalable RL & Distributed Training

Reflection AI • New York (NY)

On-site
USD 120,000 - 180,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Robotics Data Infrastructure Engineer (Distributed Systems)
Robotics Data Infrastructure Engineer (Distributed Systems)

OpenAI • Los Angeles (CA)

Hybrid
USD 230,000 - 385,000
Industrial AI & RL Engineer for Physical Systems
Industrial AI & RL Engineer for Physical Systems

FLUIX • San Francisco (CA)

On-site
USD 180,000 - 250,000
Attractive compensation package, including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth and development
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Staff Research Software Engineer - AI Training Infra
Staff Research Software Engineer - AI Training Infra

Reflection • New York (NY)

On-site
USD 120,000 - 160,000
Top-tier compensation
Health & wellness benefits
Paid parental leave
+2
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
RL Infrastructure Engineer — Scalable Training & Performance
RL Infrastructure Engineer — Scalable Training & Performance

xAI • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Health insurance
Life and AD&D insurance
Fertility benefits
+3