AI Pre-Training Engineer — Massive GPU Scale

OP Recruiting

Chicago (IL)

On-site

USD 160,000 - 260,000

Full time

43 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OP Recruiting seeks an AI Research Engineer to advance large-scale pre-training capabilities, architecting and deploying foundation models over extensive GPU infrastructure. The role involves collaboration with ML researchers to design architectures and compute roadmaps, and translating DL methods into production-grade trading models.

Applicants should have deep DL systems experience and proficiency with PyTorch, CUDA, JAX, Triton, or XLA.

Qualifications

  • 2+ years of hands-on professional experience engineering deep learning systems.
  • Deep expertise in modern hardware acceleration and deep learning frameworks (PyTorch, JAX, CUDA, Triton).
  • Proven track record of developing or tuning low-level training infrastructure or distributed training workflows.
  • Ability to solve open-ended systems performance challenges without relying on off-the-shelf software.
  • No prior finance background required.

Responsibilities

  • Optimize distributed deep learning pre-training, focusing on networking, memory, data pipelines, and fault tolerance.
  • Collaborate with ML researchers to co-design network architectures and compute roadmaps.
  • Write high-performance lower-level code and custom kernels for large-scale GPU infra.
  • Translate cutting-edge deep learning methods into production-grade trading models.

Skills

DL systems experience
High-performance computing

Tools

PyTorch
CUDA
JAX
Triton
XLA

Job description

OP Recruiting seeks an AI Research Engineer to advance large-scale pre-training capabilities, architecting and deploying foundation models over extensive GPU infrastructure. The role involves collaboration with ML researchers to design architectures and compute roadmaps, and translating DL methods into production-grade trading models.

Applicants should have deep DL systems experience and proficiency with PyTorch, CUDA, JAX, Triton, or XLA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer, Pre-Training
AI Research Engineer, Pre-Training

OP Recruiting • Chicago (IL)

On-site
USD 160,000 - 260,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior ML Engineer — Scale Training Infra & AI Deployments
Senior ML Engineer — Scale Training Infra & AI Deployments

Best AI Tools Wiki • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Equity package
Health insurance
Unlimited PTO
+3
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
AI Research Engineer (Pre-training - LLM & Multi-Modal)
AI Research Engineer (Pre-training - LLM & Multi-Modal)

Tether.io • Indiana (PA)

On-site
USD 120,000 - 160,000
AI Platform Engineer: Scale Large Models
AI Platform Engineer: Scale Large Models

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior Research Engineer - Large-Scale DL Systems (JAX/TPU)
Senior Research Engineer - Large-Scale DL Systems (JAX/TPU)

AssemblyAI • United States

Remote
USD 140,000 - 230,000
Campus AI Research Engineer: High-Impact ML for Markets
Campus AI Research Engineer: High-Impact ML for Markets

Jump Trading • Chicago (IL)

On-site
USD 270,000 - 330,000
Senior Research Engineer, Model Training / Pretraining
Senior Research Engineer, Model Training / Pretraining

Intelix.AI • Seattle (WA)

On-site
USD 147,000 - 220,000