Frontier AI Systems Engineer — Remote, Equity, Impact

Prime Intellect, Inc.

United States

Remote

USD 150,000 - 350,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work options
Visa sponsorship
Relocation assistance
Quarterly off-sites
Learning opportunities

Job summary

Prime Intellect, Inc. is building the open frontier AI stack, a full-stack platform for post-training at frontier scale. We are seeking engineers to design and optimize distributed training infrastructure and collaborate with researchers and infrastructure engineers.

You will work on kernel and runtime optimizations, multi-node GPU clusters, and open-source ML systems, with remote-friendly, SF-office options and visa sponsorship for international candidates.

Qualifications

  • Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference.
  • Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling.
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.

Responsibilities

  • Build and optimize the distributed training infrastructure behind pre-training and large-scale RL workloads.
  • Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
  • Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
  • Work on distributed training systems spanning data, tensor, and pipeline parallel workloads.
  • Shape the architecture of the RL training stack, including async rollout and post-training systems.
  • Contribute to open-source libraries and internal infrastructure used for frontier-scale model training.
  • Collaborate with researchers and infra engineers to translate bottlenecks into system improvements.
  • Stay at the frontier of training/inference systems, compiler/runtime tooling, and hardware-aware optimization.

Skills

Systems engineering
AI/ML infrastructure
PyTorch
Distributed training
CUDA / Triton kernels
Performance debugging
GPU architectures
Open-source contributions

Tools

PyTorch Distributed
DeepSpeed
FSDP
Megatron
vLLM
Ray

Job description

Prime Intellect, Inc. is building the open frontier AI stack, a full-stack platform for post-training at frontier scale. We are seeking engineers to design and optimize distributed training infrastructure and collaborate with researchers and infrastructure engineers.

You will work on kernel and runtime optimizations, multi-node GPU clusters, and open-source ML systems, with remote-friendly, SF-office options and visa sponsorship for international candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Frontier AI Training Systems Engineer
Frontier AI Training Systems Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Cash compensation and equity Incentive
Remote or SF office
Visa sponsorship and relocation
+2
Frontier AI Systems Engineer - Distributed Training (Remote)
Frontier AI Systems Engineer - Distributed Training (Remote)

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Frontier RL Systems Engineer
Frontier RL Systems Engineer

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Frontier RL Research Engineer - Remote & Visa Sponsorship
Frontier RL Research Engineer - Remote & Visa Sponsorship

Prime Intellect, Inc. • United States

Hybrid
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Quarterly team off-sites
+2
Remote Frontier AI Platform Engineer (Full Stack)
Remote Frontier AI Platform Engineer (Full Stack)

Primeintellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+2
Frontier AI Platform Engineer
Frontier AI Platform Engineer

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+3
Staff Systems Engineer, Frontier AI Infrastructure
Staff Systems Engineer, Frontier AI Infrastructure

Primeintellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Visa sponsorship
Relocation support
Professional development budget
+1
Staff Engineer, Frontier AI Compute Platform
Staff Engineer, Frontier AI Compute Platform

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 230,000
Open-Source Frontier AI Platform Engineer
Open-Source Frontier AI Platform Engineer

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
In-person SF collaboration
Visa sponsorship
Relocation assistance
+2
Remote RL Infrastructure Engineer - Frontier AI Stack
Remote RL Infrastructure Engineer - Frontier AI Stack

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites