ML Infrastructure Engineer: Training, Inference & Kernels

Trajectory

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Trajectory is seeking an ML Infrastructure Engineer to build the platform that powers training, inference, and kernels for next‑gen AI systems. You will own distributed training, low‑latency serving, and GPU‑kernel optimizations, shaping reusable systems that scale in production.

You will collaborate with researchers to turn new algorithms into reliable components, build benchmarks and observability, and push measurable gains while preserving correctness.

Qualifications

  • Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping and operating demanding systems.
  • Deep specialization in training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
  • Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
  • Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
  • Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.
  • We value demonstrated capability over credentials.

Responsibilities

  • Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles.
  • Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
  • Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.
  • Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements. Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.
  • We’re building one of the world’s best continual learning loops for ML infrastructure: agents propose optimizations, run experiments, measure gains, and learn from the results across training, inference, and kernels.

Skills

Distributed systems
Networking
Storage
Failure recovery
GPU kernels
Python
CUDA
C++/CUDA/Triton
Profiling & debugging
Ownership & communication

Tools

CUDA
Triton
PyTorch

Job description

Trajectory is seeking an ML Infrastructure Engineer to build the platform that powers training, inference, and kernels for next‑gen AI systems. You will own distributed training, low‑latency serving, and GPU‑kernel optimizations, shaping reusable systems that scale in production.

You will collaborate with researchers to turn new algorithms into reliable components, build benchmarks and observability, and push measurable gains while preserving correctness.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Trajectory • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infrastructure Engineer: Training & Inference
ML Infrastructure Engineer: Training & Inference

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
ML Systems Engineer: Trainium Inference & Kernels
ML Systems Engineer: Trainium Inference & Kernels

Slope • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infrastructure Engineer: Training & Real-Time Inference
ML Infrastructure Engineer: Training & Real-Time Inference

Mach9 Robotics Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000