Frontier AI Training Systems Engineer

Prime Intellect AI

San Francisco (CA)

Hybrid

USD 150,000 - 350,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Cash compensation and equity Incentive
Remote or SF office
Visa sponsorship and relocation
Team off-sites and learning
Open-source collaboration

Job summary

Prime Intellect AI is building the open superintelligence stack and frontier AI infrastructure. You will help design, implement, and optimize distributed training systems for pre-training and RL workloads, collaborating with researchers and engineers to push the boundaries of model-scale and performance.

You’ll work on kernels, runtimes, and multi-node GPU setups, contributing to open-source ML systems and tooling while enjoying flexible work arrangements—remote or in-person in San

Qualifications

  • Strong systems engineering experience in AI/ML infrastructure.
  • Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling.
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong understanding of GPU architecture, profiling, and performance debugging.
  • Ability to identify bottlenecks across the stack and drive improvements from first principles.
  • Comfort working in a fast-moving environment with ambiguous problems and high ownership.

Responsibilities

  • Build and optimize the distributed training infrastructure behind pre-training and large-scale RL training workloads by contributing to the prime-rl framework.
  • Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
  • Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
  • Work on distributed training systems spanning data, tensor, and pipeline parallel workloads.
  • Help shape the architecture of the RL training stack, including async rollout and post-training systems.
  • Contribute to open-source libraries and internal infrastructure used for frontier-scale model training.
  • Collaborate closely with researchers and infrastructure engineers to translate bottlenecks into concrete systems improvements.
  • Stay at the frontier of training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.

Job description

Prime Intellect AI is building the open superintelligence stack and frontier AI infrastructure. You will help design, implement, and optimize distributed training systems for pre-training and RL workloads, collaborating with researchers and engineers to push the boundaries of model-scale and performance.

You’ll work on kernels, runtimes, and multi-node GPU setups, contributing to open-source ML systems and tooling while enjoying flexible work arrangements—remote or in-person in San

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier AI Systems Engineer - Distributed Training (Remote)
Frontier AI Systems Engineer - Distributed Training (Remote)

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Frontier RL Systems Engineer
Frontier RL Systems Engineer

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Remote RL Infrastructure Engineer - Frontier AI Stack
Remote RL Infrastructure Engineer - Frontier AI Stack

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites
Frontier AI Platform Engineer
Frontier AI Platform Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+3
Senior GPU Infrastructure Architect for Frontier AI
Senior GPU Infrastructure Architect for Frontier AI

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Equity incentives
Staff Engineer, Frontier AI Infrastructure
Staff Engineer, Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range: $150-300k
Flexible work arrangement (SF office +
Full visa sponsorship and relocation
+1
RL Research Engineer for Frontier AI Platform
RL Research Engineer for Frontier AI Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote work option
Visa sponsorship
Relocation assistance
+2
Staff Engineer – Open Frontier AI Infrastructure
Staff Engineer – Open Frontier AI Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Staff Storage Engineer – Frontier AI Infra
Staff Storage Engineer – Frontier AI Infra

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Frontier AI Storage Infrastructure Engineer
Frontier AI Storage Infrastructure Engineer

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000