Senior ML Infra Engineer for Large-Scale Mid-Training & RL

Peano AI

Palo Alto (CA)

On-site

USD 200,000 - 260,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Peano AI is seeking a Founding Large-Scale Mid-Training/RL Infrastructure Engineer to design and scale our in-house training stack for massive language models. You will work with thousands of GPUs/TPUs, optimize pretraining and RL pipelines, and push the performance envelope for end-to-end training workflows.

The role is deeply technical with ownership across architecture, data pipelines, sharding, and deployment.

Qualifications

  • Extensive experience with large-scale DL training and distributed systems.
  • Proven record in pretraining or RL-based fine-tuning of large models.
  • Hands-on with Megatron, Transformer-Engine, verl, slime or similar toolkits.
  • Experience designing data pipelines, sharding, checkpointing, and training ops at scale.

Responsibilities

  • Build, optimize, and scale distributed training infrastructure for foundation models.
  • Improve throughput, efficiency, scalability, cost, and reliability of pretraining and RL pipelines.
  • Architect data loading, sharding, checkpointing, batching, and accelerator use for multi-node training.
  • Integrate and optimize RL components: rollout, reward modeling, and environment orchestration.
  • Work with Megatron, Transformer-Engine, verl, slime, and related toolkits.
  • Tune memory usage, mixed-precision, and runtime performance for massive models.
  • Debug and profile training bottlenecks across code, compute, networking, and infra.
  • Collaborate with teams to ensure performance and functionality of foundation models.
  • Translate research prototypes into production-ready, scalable systems.

Skills

Distributed systems design
Large-scale DL training
Performance tuning
Debugging & profiling
Strong ownership

Tools

Megatron
Transformer-Engine
verl
slime

Job description

Peano AI is seeking a Founding Large-Scale Mid-Training/RL Infrastructure Engineer to design and scale our in-house training stack for massive language models. You will work with thousands of GPUs/TPUs, optimize pretraining and RL pipelines, and push the performance envelope for end-to-end training workflows.

The role is deeply technical with ownership across architecture, data pipelines, sharding, and deployment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Large-Scale RL & Mid-Training Infra Engineer
Lead Large-Scale RL & Mid-Training Infra Engineer

Engg • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Founding Mid-Training/RL Infrastructure Engineer
Founding Mid-Training/RL Infrastructure Engineer

Engg • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Founding ML Infra Engineer — Equity & Open-Source AI
Founding ML Infra Engineer — Equity & Open-Source AI

Peano AI • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Founding Mid-Training/RL Infrastructure Engineer
Founding Mid-Training/RL Infrastructure Engineer

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Senior ML Engineer — Scale Training Infra & AI Deployments
Senior ML Engineer — Scale Training Infra & AI Deployments

Best AI Tools Wiki • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Equity package
Health insurance
Unlimited PTO
+3
RL Infrastructure Engineer: Scale End-to-End ML
RL Infrastructure Engineer: Scale End-to-End ML

Lever, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Medical benefits
Dental benefits
Vision benefits
+7
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1