Member of Technical Staff - ML Infra

Trajectory

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Trajectory is seeking an ML Infrastructure Engineer to build the platform that powers training, inference, and kernels for next‑gen AI systems. You will own distributed training, low‑latency serving, and GPU‑kernel optimizations, shaping reusable systems that scale in production.

You will collaborate with researchers to turn new algorithms into reliable components, build benchmarks and observability, and push measurable gains while preserving correctness.

Qualifications

  • Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping and operating demanding systems.
  • Deep specialization in training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
  • Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
  • Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
  • Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.
  • We value demonstrated capability over credentials.

Responsibilities

  • Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles.
  • Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
  • Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.
  • Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements. Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.
  • We’re building one of the world’s best continual learning loops for ML infrastructure: agents propose optimizations, run experiments, measure gains, and learn from the results across training, inference, and kernels.

Skills

Distributed systems
Networking
Storage
Failure recovery
GPU kernels
Python
CUDA
C++/CUDA/Triton
Profiling & debugging
Ownership & communication

Tools

CUDA
Triton
PyTorch

Job description

Job Description

As an ML Infrastructure Engineer at Trajectory, you will build the infrastructure for AI systems that learn to improve their own training, inference, and kernels.

Our goal is to unlock a 10× improvement every month somewhere in the stack - from GPU scale and model size to throughput, memory efficiency, caching, and latency.

This role spans training, inference, and kernels. Bring deep expertise in at least one area and curiosity across the stack. We’ll shape your initial ownership around your strengths.

What you will work on
Training

Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles.

Inference

Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.

Kernels and runtimes

Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.

Across all three

Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements. Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.

Continual learning for ML infrastructure

We’re building one of the world’s best continual learning loops for ML infrastructure: agents propose optimizations, run experiments, measure gains, and learn from the results across training, inference, and kernels.

Required Qualifications
  • Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping and operating demanding systems.
  • Required: deep specialization in training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
  • Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
  • Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
  • Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.

We value demonstrated capability over credentials.

Preferred

Hands-on experience with training and RL stacks such as Miles, SkyRL, Prime Intellect’s verifiers, or comparable systems. Depending on your specialty, experience with vLLM, SGLang, collective communication, or ML compilers is also valuable.

About Trajectory

Trajectory is a research and product lab creating the platform for continual learning.

AI is the most capable software ever built, and the least able to learn. Every valuable correction and edit that happens in a product evaporates at the next session. A few teams have closed this gap by hand-coupling their models to their products: Composer, Claude Code, Windsurf SWE-1.

Trajectory is the first scalable approach for every company: our platform unlocks the signal already sitting in product use, so companies can continuously post-train large-scale agentic models that outperform the frontier.

Our research team comes from Deepmind, OpenAI, Meta Superintelligence, and product team from Figma, Apple, Stripe, and Windsurf. We’re working with customers like Harvey, Rogo, Mercor, Decagon and Clay, and we’ve raised $60M from Sequoia, Conviction, Jeff Dean, and Fei Fei Li.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Generalist
Member of Technical Staff - Generalist

Trajectory • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff - Design Eng
Member of Technical Staff - Design Eng

Trajectory • San Francisco (CA)

On-site
USD 120,000 - 150,000
Member of Business Staff
Member of Business Staff

Trajectory • San Francisco (CA)

On-site
USD 140,000 - 190,000
ML Infrastructure Engineer: Training, Inference & Kernels
ML Infrastructure Engineer: Training, Inference & Kernels

Trajectory • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Observable Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000
Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Collective Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

On-site
USD 140,000 - 210,000
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2