Member of Technical Staff — Training Infrastructure

Kindredventures

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress.

Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling.

Qualifications

  • Experience with distributed training frameworks and techniques to train large foundation models.
  • Strong grasp of parallelism, memory optimization, mixed precision, and communication overlap.
  • Ability to profile and debug performance in complex codebases from framework internals down to kernels and collectives.
  • Deep understanding of PyTorch and JAX and their underlying system architectures.
  • Bonus: contributions to open-source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA).

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs
  • Research and test parallelization strategies and numerical precision trade-offs across model scales
  • Analyze, profile, and debug low-level GPU operations to maximize throughput and hardware utilization
  • Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration
  • Collaborate with researchers to bring novel model architectures from prototype to full scale
  • Stay up-to-date on research to bring new ideas to work

Skills

Distributed training
FSDP
DeepSpeed
Megatron
PyTorch
JAX/XLA
Performance profiling
Memory optimization
Mixed precision
Kernels/Collectives
Open-source contributions

Tools

N/A

Job description

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.

To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.

Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.

We look for infrastructure engineers who are excited to tackle unsolved problems. Training an LPM means scaling novel architectures over multimodal physical data — a problem where the playbooks from language and vision only partially apply. Your mission is to make large‑scale training fast, efficient, and reliable, so that every GPU cycle accelerates research progress.

Responsibilities
  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs
  • Research and test parallelization strategies and numerical precision trade-offs across model scales, including for architectures that don't map cleanly onto existing LLM training stacks
  • Analyze, profile, and debug low-level GPU operations to maximize throughput and hardware utilization
  • Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration
  • Collaborate with researchers to bring novel model architectures from prototype to full scale
  • Stay up-to-date on research to bring new ideas to work
What we’re looking for

We value a relentless approach to problem‑solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Demonstrated proficiency with distributed training frameworks and techniques (e.g. FSDP, DeepSpeed, Megatron, Pytorch, JAX/XLA) to train large foundation models
  • Strong grasp of state‑of‑the‑art techniques for optimizing training workloads: parallelism strategies, memory optimization, mixed precision, communication overlap
  • Ability to profile and debug performance in complex codebases, from framework internals down to kernels and collectives
  • Deep understanding of deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures
  • Bonus: contributions to open‑source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Research, Physics
Member of Technical Staff — Research, Physics

Kindredventures • San Francisco (CA)

On-site
USD 150,000 - 230,000