Distributed Systems Engineer - Data & Inference Platform

OpenTalent

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible work
Adaption Passport
Lunch Stipend
Well-Being

Job summary

OpenTalent in San Francisco is seeking a systems-minded engineer to build and operate inference services for large language models at scale. You will own distributed pipelines and GPU memory optimization, collaborating with researchers and ML engineers to deliver reliable, cost-efficient systems.

The role emphasizes production reliability, incident ownership, and Bay Area in-person collaboration, with opportunities to shape the stack from data ingest to model serving.

Qualifications

  • 5+ years building and operating distributed systems in production.
  • Deep experience with at least one large-scale data or compute framework (Ray, Spark, Flink, Beam, Dask).
  • Strong fluency in Python and at least one systems language (Go, Rust, C++).
  • Working knowledge of the GPU/accelerator stack: CUDA fundamentals, NCCL, mixed precision, memory layout.
  • Experience operating Kubernetes-based infrastructure, including custom operators or schedulers.
  • A track record of owning hard production incidents end-to-end — diagnosis, mitigation, and the durable fix.
  • Bonus: hands-on experience with LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI), modern lakehouse formats (Iceberg, Delta, Hudi), or open-source contributions to relevant projects.

Responsibilities

  • Serve Models at Scale: Design and operate distributed inference systems for LLMs, optimizing throughput, latency, and cost across heterogeneous GPU fleets. Batching, scheduling, KV cache management, autoscaling — you own the levers that make inference economical.
  • Move the Data: Build large-scale data pipelines (Ray Data, Spark, or equivalents) that ingest, transform, and curate the datasets behind training and evaluation. The bottleneck is rarely where people think it is, and you find it.
  • Debug the Undebuggable: Chase down the failure modes that only emerge under real production traffic — stragglers, head-of-line blocking, silent data corruption, GPU memory fragmentation — and write the postmortems that prevent the next ten. Define SLOs, build the observability to measure them, and own the on-call rotation that defends them.
  • Partner Across the Stack: Work directly with researchers and ML engineers to take experimental workloads from "runs on one node" to "runs in production." You're a systems partner, not a ticket queue.

Tools

Ray
Spark
Flink
Kubernetes
CUDA
Python
Go
Rust
C++

Job description

OpenTalent in San Francisco is seeking a systems-minded engineer to build and operate inference services for large language models at scale. You will own distributed pipelines and GPU memory optimization, collaborating with researchers and ML engineers to deliver reliable, cost-efficient systems.

The role emphasizes production reliability, incident ownership, and Bay Area in-person collaboration, with opportunities to shape the stack from data ingest to model serving.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud-Scale Backend Engineer for ML Inference
Cloud-Scale Backend Engineer for ML Inference

Praxis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 250,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Staff Software Engineer — ML Platform & Inference
Staff Software Engineer — ML Platform & Inference

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 280,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Inference Platform Backend Engineer (Equity & Benefits)
Inference Platform Backend Engineer (Equity & Benefits)

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity