Inference Infrastructure Engineer for Large-Scale AI

Causal

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Causal is building a Large Physics foundation Model and seeks infrastructure engineers for high-throughput, low-latency inference at scale in San Francisco. You will work on evaluating physical observations, backtesting, and large-batch workflows, collaborating with researchers to push model performance.

We value deep learning expertise, GPU-aware optimization, and production-grade engineering. Proficiency with PyTorch or JAX, Kubernetes-based orchestration, and open‑source inference tooling is

Qualifications

  • Experience building or optimizing inference and serving systems for throughput and latency (e.g. TensorRT).
  • Understanding of distributed compute, GPU parallelism, and hardware‑aware optimization.
  • Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures.
  • Strong engineering skills: performant, maintainable code and the ability to debug complex codebases.

Responsibilities

  • Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations.
  • Design and implement techniques that improve latency, throughput, and efficiency for real-time inference.
  • Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory.
  • Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps.
  • Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable.
  • Collaborate with researchers to enable high-performance inference for novel architectures as they emerge.

Skills

Inference systems
Latency optimization
Distributed compute
GPU parallelism
PyTorch
JAX
Deep learning frameworks
Kubernetes
TensorRT

Tools

Ray
Slurm

Job description

Causal is building a Large Physics foundation Model and seeks infrastructure engineers for high-throughput, low-latency inference at scale in San Francisco. You will work on evaluating physical observations, backtesting, and large-batch workflows, collaborating with researchers to push model performance.

We value deep learning expertise, GPU-aware optimization, and production-grade engineering. Proficiency with PyTorch or JAX, Kubernetes-based orchestration, and open‑source inference tooling is

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer - Large-Scale AI Models & Pipelines
Research Engineer - Large-Scale AI Models & Pipelines

Magic • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer, AI Production & Customer Platforms
Staff Engineer, AI Production & Customer Platforms

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 190,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6