Member of Technical Staff — Inference Infrastructure

Causal Labs

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations.

You will design techniques to improve latency and throughput, optimize the inference stack to exhaust hardware, and extend Kubernetes, Ray, and Slurm for distributed inference, ensuring reliability and reproducibility across the stack in collaboration with researchers.

Qualifications

  • Experience building inference/serving systems for throughput and latency.
  • Understanding of distributed compute, GPU parallelism, and hardware-aware optimization.
  • Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures.
  • Strong engineering skills: performant, maintainable code and the ability to debug complex codebases.
  • Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton).

Responsibilities

  • Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations
  • Design and implement techniques that improve latency, throughput, and efficiency for real-time inference
  • Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory
  • Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps
  • Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable
  • Collaborate with researchers to enable high-performance inference for novel architectures as they emerge

Skills

High-throughput inference systems
Distributed compute
GPU acceleration
PyTorch/JAX
Kubernetes/Ray/Slurm
Performance debugging

Tools

TensorRT
CUDA

Job description

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.

To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.

Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.

We look for infrastructure engineers who are excited to tackle unsolved problems. Progress on an LPM is gated by how fast we can evaluate it: large-scale backtesting against decades of physical observations, ensemble generation, and rollout evaluation across model scales.

Responsibilities

  • Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations

  • Design and implement techniques that improve latency, throughput, and efficiency for real-time inference

  • Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory

  • Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps

  • Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable

  • Collaborate with researchers to enable high-performance inference for novel architectures as they emerge

What we\'re looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Experience building or optimizing inference and serving systems for throughput and latency (e.g. TensorRT)

  • Understanding of distributed compute, GPU parallelism, and hardware-aware optimization

  • Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures

  • Strong engineering skills: performant, maintainable code and the ability to debug complex codebases

  • Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Product Engineering
Member of Technical Staff — Product Engineering

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000