Staff Inference Systems Engineer — High-Throughput AI

Kindredventures

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm.

Ideal candidates have a track record in building efficient inference stacks, GPU-aware optimization, and deep learning frameworks like PyTorch or JAX. Collaboration with researchers and a focus on reliability are essential.

Qualifications

  • Experience building or optimizing inference/serving systems for throughput and latency.
  • Understanding of distributed compute, GPU parallelism, and hardware-aware optimization.
  • Deep familiarity with deep learning frameworks (PyTorch, JAX) and their underlying system architectures.
  • Strong engineering skills: performant, maintainable code and the ability to debug complex codebases.
  • Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, Triton).

Responsibilities

  • Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical observations.
  • Design and implement techniques to improve latency, throughput, and efficiency for real-time inference.
  • Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory.
  • Extend orchestration frameworks (Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps.
  • Establish standards for reliability, observability, and reproducibility across the inference stack.
  • Collaborate with researchers to enable high-performance inference for novel architectures as they emerge.

Skills

Inference systems
Distributed compute
GPU optimization
PyTorch
JAX
Maintainable code

Tools

TensorRT
Kubernetes
Ray
Slurm

Job description

Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm.

Ideal candidates have a track record in building efficient inference stacks, GPU-aware optimization, and deep learning frameworks like PyTorch or JAX. Collaboration with researchers and a focus on reliability are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Inference Infrastructure Engineer for Large-Scale AI
Inference Infrastructure Engineer for Large-Scale AI

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Systems Engineer — Inference Runtime Lead
Staff Systems Engineer — Inference Runtime Lead

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 405,000 - 485,000
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

causal • San Francisco (CA)

On-site
USD 190,000 - 270,000
Staff AI Infrastructure Engineer — Orchestration & Inference
Staff AI Infrastructure Engineer — Orchestration & Inference

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership