Remote RL Infrastructure Engineer - Frontier AI Stack

Prime Intellect AI

San Francisco (CA)

Hybrid

USD 150,000 - 350,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites

Job summary

Prime Intellect AI is building an open frontier AI stack that enables frontier-scale model training and deployment. The role focuses on designing and optimizing the underlying systems infrastructure for large-scale RL and distributed training workloads.

You will work closely with researchers and infrastructure engineers to push the limits of training performance, kernels, and runtime optimizations, while contributing to open-source projects and internal tooling for frontier-scale models.

Qualifications

  • Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference.
  • Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling.
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong understanding of GPU architecture, profiling, and performance debugging.
  • Ability to identify bottlenecks across the stack and drive improvements from first principles.
  • Comfort working in a fast-moving environment with ambiguous problems and high ownership.

Responsibilities

  • Build and optimize the systems infrastructure behind large-scale RL and distributed training workloads by contributing to our prime-rl framework.
  • Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
  • Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
  • Work on distributed training systems spanning data, tensor, and pipeline parallel workloads.
  • Help shape the architecture of our RL training stack, including async rollout and post-training systems.
  • Contribute to open-source libraries and internal infrastructure used for frontier-scale model training.
  • Collaborate closely with researchers and infrastructure engineers to translate bottlenecks into concrete systems improvements.
  • Stay at the frontier of training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.

Skills

Systems engineering
AI/ML infrastructure
PyTorch
Distributed training
GPU architecture

Tools

DeepSpeed
Megatron
vLLM
Ray

Job description

Prime Intellect AI is building an open frontier AI stack that enables frontier-scale model training and deployment. The role focuses on designing and optimizing the underlying systems infrastructure for large-scale RL and distributed training workloads.

You will work closely with researchers and infrastructure engineers to push the limits of training performance, kernels, and runtime optimizations, while contributing to open-source projects and internal tooling for frontier-scale models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier RL Systems Engineer
Frontier RL Systems Engineer

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
RL Research Engineer for Frontier AI Platform
RL Research Engineer for Frontier AI Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote work option
Visa sponsorship
Relocation assistance
+2
Frontier AI Training Systems Engineer
Frontier AI Training Systems Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Cash compensation and equity Incentive
Remote or SF office
Visa sponsorship and relocation
+2
Frontier AI Systems Engineer - Distributed Training (Remote)
Frontier AI Systems Engineer - Distributed Training (Remote)

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Staff Engineer – Open Frontier AI Infrastructure
Staff Engineer – Open Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Staff Engineer, Frontier AI Infrastructure
Staff Engineer, Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range: $150-300k
Flexible work arrangement (SF office +
Full visa sponsorship and relocation
+1
Frontier AI Platform Engineer
Frontier AI Platform Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+3
Lead Security Engineer for Frontier AI Infrastructure
Lead Security Engineer for Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 180,000 - 350,000
Remote or SF office option
Visa sponsorship
Professional development budget
+2
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites