ML Systems Engineer — Scalable Distributed Training & Inference

OpenTalent

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Sciforium is seeking a distributed training and inference engineer to build and optimize the ML software stack for large-scale AI workloads. You will work across CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch to ensure fast, scalable training and serving.

This role emphasizes deep systems engineering, debugging hardware–software interactions, and optimizing performance at every layer of the ML stack, enabling training and deployment of next‑gen LLMs and generative AI models.

Qualifications

  • 5+ years of industry experience in ML systems or distributed training.
  • Strong programming experience in Python and C++ with distributed systems.
  • Familiarity with ML tooling and profiling tools.
  • Solid academic background in CS/CE/EE.

Responsibilities

  • Maintain and optimize ML libraries and frameworks across environments.
  • Own end-to-end ML software stack from drivers to tooling.
  • Tune distributed training for scalability and efficiency.
  • Profile and optimize performance across multi-node clusters.
  • Debug hardware–software interactions and ensure stability.
  • Collaborate with research and kernel teams to improve throughput.

Skills

Python
C++
Distributed systems
ML tooling

Education

Bachelor's or Master’s degree in CS/CE/EE

Tools

Nsight
ROCm Profiler
XLA profiler
TPU tools

Job description

Sciforium is seeking a distributed training and inference engineer to build and optimize the ML software stack for large-scale AI workloads. You will work across CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch to ensure fast, scalable training and serving.

This role emphasizes deep systems engineering, debugging hardware–software interactions, and optimizing performance at every layer of the ML stack, enabling training and deployment of next‑gen LLMs and generative AI models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Training and Inference Engineer - sciforium
Distributed Training and Inference Engineer - sciforium

OpenTalent • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Training Infra Engineer for Scalable LLM Systems
Training Infra Engineer for Scalable LLM Systems

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4