Senior Inference Systems Engineer for High-Performance AI

Causal Labs

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations.

You will design techniques to improve latency and throughput, optimize the inference stack to exhaust hardware, and extend Kubernetes, Ray, and Slurm for distributed inference, ensuring reliability and reproducibility across the stack in collaboration with researchers.

Qualifications

  • Experience building inference/serving systems for throughput and latency.
  • Understanding of distributed compute, GPU parallelism, and hardware-aware optimization.
  • Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures.
  • Strong engineering skills: performant, maintainable code and the ability to debug complex codebases.
  • Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton).

Responsibilities

  • Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations
  • Design and implement techniques that improve latency, throughput, and efficiency for real-time inference
  • Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory
  • Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps
  • Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable
  • Collaborate with researchers to enable high-performance inference for novel architectures as they emerge

Skills

High-throughput inference systems
Distributed compute
GPU acceleration
PyTorch/JAX
Kubernetes/Ray/Slurm
Performance debugging

Tools

TensorRT
CUDA

Job description

Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations.

You will design techniques to improve latency and throughput, optimize the inference stack to exhaust hardware, and extend Kubernetes, Ray, and Slurm for distributed inference, ensuring reliability and reproducibility across the stack in collaboration with researchers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Infrastructure Engineer for Large-Scale AI
Inference Infrastructure Engineer for Large-Scale AI

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Systems Engineer, High-Throughput AI Serving
Inference Systems Engineer, High-Throughput AI Serving

Future Ventures • Palo Alto (CA)

On-site
USD 135,000 - 160,000
Comprehensive medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
+1
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, Scaling Inference Systems
Staff Software Engineer, Scaling Inference Systems

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000