Founding Mid-Training/RL Infrastructure Engineer

Engg

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Peano AI is seeking a Founding Large-Scale Mid-Training/RL Infrastructure Engineer to design, run, and scale our in-house training stack. You will work with thousands of GPUs and TPU accelerators to optimize pretraining and RL-based methods for large language models, impacting core products and future roadmap.

This role is deeply technical, with responsibilities spanning data pipelines, sharding, checkpointing, and accelerator-aware scheduling.

Qualifications

  • Experience with large-scale deep learning model training and distributed systems.
  • Proven track record with pretraining and RL-based fine-tuning of large models.
  • Hands-on with Megatron, Transformer-Engine, verl, slime.
  • Experience in data pipelines, model/data/optimizer sharding, and checkpointing.
  • Strong accelerator/memory optimization skills at cluster scale.
  • Excellent debugging and performance-tuning in distributed environments.
  • Comfort working in deeply technical, early-stage startup settings.

Responsibilities

  • Build, optimize, and scale distributed training infrastructure for foundation models.
  • Own throughput, efficiency, scalability, cost, and reliability of pretraining and RL pipelines.
  • Architect data loading, sharding, checkpointing, batching for multi-node training.
  • Integrate and optimize RL components: rollout, reward modeling, environment orchestration.
  • Work with Megatron, Transformer-Engine, verl, slime and other libraries.
  • Tune memory, mixed-precision, and runtime performance for massive models.
  • Debug and profile training bottlenecks across code, compute, networking, and infra.
  • Collaborate with research, infra, and applications teams to deliver performant models.

Skills

Large-scale model training
Distributed systems
RL training
Performance tuning
Debugging
Startup mindset

Tools

Megatron
Transformer-Engine
verl
slime

Job description

Founding Large-Scale Mid-Training/RL Infrastructure Engineer Location: Onsite in Palo Alto Compensation: Competitive Salary + Equity

ABOUT PEANO AI

Peano AI is building the infrastructure and application stack for the next generation of agentic AI systems. We believe token usage will grow exponentially over the coming years, but routing all inference and training through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving and training infrastructure paired with an application layer designed for long-running, agentic workloads. Peano AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, large-scale open-source model deployment, and pretraining of our own foundation model. By combining infrastructure and application design, we aim to make open-source models and custom foundation models significantly more performant, practical, and competitive.

ABOUT THIS ROLE

We are looking for a Large-Scale Mid-Training/RL Infrastructure Engineer to help build, optimize, and scale out our in-house foundation model training stack, with an emphasis on both core pretraining and RL-enhanced methods. This role is deeply technical and directly impacts Peano AI's core model product. You will have a rare opportunity to design and run distributed infrastructure at massive scale using thousands of the latest NVIDIA GB300, VR200 GPUs, and next-generation TPU v7x accelerators. You'll tackle ambitious optimization and scaling challenges with state‑of‑the‑? You will architect, scale, and tune the distributed and accelerator‑aware training infrastructure for massive language models, including data pipelines, model sharding, checkpointing, rollout and reward‑based RL, and overall system throughput. Experience with large‑scale model pretraining and mid‑training interventions is required. Preferred experience with frameworks such as Megatron, Transformer‑Engine, verl, slime, and related large‑scale training and RL toolkits. Familiarity with training pipeline scaling, mixed‑precision, memory optimization, and deep understanding of both supervised and RL‑based training cycles is essential.

WHAT YOU’LL DO
  • Build, optimize, and scale the distributed training infrastructure for Peano AI’s foundation models, leveraging thousands of NVIDIA GB300/VR200 GPUs and TPU v7x accelerators for high throughput and efficiency.
  • Own and improve throughput, efficiency, scalability, cost, and reliability of end‑to‑end pretraining and RL pipelines.
  • Architect distributed data loading, sharding, checkpointing, batching, and accelerator utilization for multi‑node LLM training at unmatched scale.
  • Integrate and optimize RL components: rollout, reward modeling, environment orchestration, and mid‑training signal injections.
  • Work with and extend frameworks such as Megatron, Transformer‑Engine, verl, slime, and other high‑performance training libraries.
  • Tune memory usage, mixed‑precision, and runtime performance for massive models and long‑context workloads.
  • Debug and profile training performance bottlenecks—across model code, distributed compute, networking, and infrastructure.
  • Collaborate with application, infrastructure, and research teams to ensure our foundation models deliver on performance and functionality.
  • Translate research prototypes and experimental features into production‑ready, scalable training systems that operate across state‑of‑the‑?
QUALIFICATIONS
  • Significant experience with large-scale deep learning model training and distributed system design, including large GPU/TPU clusters (GB300/VR200, TPU v7x or similar).
  • Proven track record of pretraining and/or RL‑based fine‑tuning of large models (LLMs or comparable scale).
  • Deep familiarity and hands‑on experience with frameworks such as Megatron, Transformer‑Engine, verl, slime, or similar large‑scale/foundation-model toolkits.
  • Experience with data pipeline design, model/data/optimizer sharding, checkpointing, rollout/reward design, and training operations at scale.
  • Strong accelerator (NVIDIA GB300/VR200 GPUs, TPU v7x) and memory optimization skills, including at cluster scale.
  • Excellent debugging, profiling, and performance‑tuning abilities in distributed environments.
  • Comfort working in deeply technical, high‑ownership, early‑stage startup settings.
CULTURAL FIT
  • Hands‑on technical excellence and strong engineering judgment
  • End‑to‑end ownership—from design to implementation to production
  • Bias for action: ship quickly, learn from failures, iterate
  • High intensity during critical milestones, focused on customer and product outcomes
  • Ability to work deeply and sustain high execution
  • Clear communicator, low ego, strong collaborator
  • Thrives in ambiguity, rapid change, and taking on multiple roles
  • Lifelong learner with a belief that capability compounds with time and effort

If you are excited to build the core training and RL infrastructure for Peano AI’s vertically integrated foundation model—leveraging thousands of NVIDIA GB300/VR200 GPUs and TPU v7x accelerators, tackling the hardest scaling and optimization challenges, and helping push the boundaries of agentic AI—we’d love to talk.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Mid-Training/RL Infrastructure Engineer
Founding Mid-Training/RL Infrastructure Engineer

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Founding Machine Learning Infrastructure Engineer
Founding Machine Learning Infrastructure Engineer

Peano AI • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 350,000
Remote or SF office work option
Visa sponsorship & relocation
Quarterly team offsites
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Research Engineer - RL Infrastructure
Research Engineer - RL Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Visa sponsorship
Relocation assistance
Remote work option
Lead Large-Scale RL & Mid-Training Infra Engineer
Lead Large-Scale RL & Mid-Training Infra Engineer

Engg • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Remote or SF office
Visa sponsorship
Relocation support
+3
Senior ML Infra Engineer for Large-Scale Mid-Training & RL
Senior ML Infra Engineer for Large-Scale Mid-Training & RL

Peano AI • Palo Alto (CA)

On-site
USD 200,000 - 260,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff AI Training Infrastructure Engineer
Staff AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays