Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fleet AI, Inc. is seeking a Member of Technical Staff, Research Engineer, Infrastructure to own and optimize the large GPU training cluster and the end-to-end research stack in a fast-moving environment.

You will profile bottlenecks, implement kernel and system-level fixes, and collaborate with researchers to ensure reliable, scalable training and inference at scale onsite in San Francisco.

Qualifications

  • Experience with multi-node training and major parallelism strategies (DP, TP, PP, FSDP).
  • Ability to profile slow training runs and implement fixes in CUDA/Triton kernels or data pipelines.
  • Experience building or maintaining high-throughput inference stacks (vLLM, SGLang) and KV cache tuning.
  • Ability to navigate large, evolving research codebases and improve performance without breaking correctness.

Responsibilities

  • Own the training cluster: capacity, failover, fault tolerance, on-call when training stalls.
  • Profile training runs to identify bottlenecks (NCCL, KV cache, dataloader, allgather) and implement fixes; write kernels when needed.
  • Maintain training and inference pipelines: distributed training (FSDP, TP, PP), rollout generation, async training, inference stack dependencies.
  • Maintain the research codebase: keep it fast and reliable as researchers extend it without breaking the science.

Skills

Distributed training
Kernel optimization
Inference systems
Research-codebase maintenance

Tools

CUDA
Triton

Job description

Fleet AI, Inc. is seeking a Member of Technical Staff, Research Engineer, Infrastructure to own and optimize the large GPU training cluster and the end-to-end research stack in a fast-moving environment.

You will profile bottlenecks, implement kernel and system-level fixes, and collaborate with researchers to ensure reliable, scalable training and inference at scale onsite in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Staff Research Engineer - Scalable ML Training Systems
Staff Research Engineer - Scalable ML Training Systems

Fireworks AI • San Mateo (CA)

On-site
USD 250,000 - 400,000
Research Engineer: Large-Scale AI Training Systems
Research Engineer: Large-Scale AI Training Systems

BlackForestLabs • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
Equity
Infrastructure Kernel Engineer for Scalable AI Training
Infrastructure Kernel Engineer for Scalable AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior ML Training Infrastructure Engineer
Senior ML Training Infrastructure Engineer

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Staff Research Engineer: Scalable ML & Physics Infra
Staff Research Engineer: Scalable ML & Physics Infra

Socket.dev • Cambridge (MA)

On-site
USD 224,000 - 294,000
Equity
Benefits package
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000